Please use this identifier to cite or link to this item: http://dx.doi.org/10.25673/115107
Full metadata record
DC FieldValueLanguage
dc.contributor.authorMendikowski, Melle-
dc.contributor.authorSchindler, Benjamin-
dc.contributor.authorSchmid, Thomas-
dc.contributor.authorMöller, Ralf-
dc.contributor.authorHartwig, Mattis-
dc.date.accessioned2024-03-04T08:47:26Z-
dc.date.available2024-03-04T08:47:26Z-
dc.date.issued2023-
dc.identifier.urihttps://opendata.uni-halle.de//handle/1981185920/117063-
dc.identifier.urihttp://dx.doi.org/10.25673/115107-
dc.description.abstractConsidering the growing global demand for machine learning training data, synthetic data generation is a reasonable way to address the versatile challenges in data acquisition. Conditional Tabular Generative Adversarial Network (CTGAN), an extension of the widely used Generative Adversarial Network (GAN), is considered one of the most promising techniques in the field of tabular data generation. Despite numerous successes of CTGAN, a lack of preserving categorical dependencies within the data has been identified. In prior work, the Cramer’s V (CV) as a natural metric for representing the correlation of categorical dependencies was proposed for hyperparameter tuning of CTGAN models. In this paper, we explore two novel strategies to directly integrate CV statistics of data batches within CTGAN training. The first approach is a generator loss term that penalizes differences between the CV statistics of the original and generated data. The second innovation is the extraction of the CV matrix as an additional feature for the critic. By applying our proposed methods to three benchmark datasets, we improve the averaged accuracy of supervised learning models trained on synthesized data by 11 % compared to the legacy CTGAN. We also outline the impact of CV statistics on preserving dependencies between categorical data columns in terms of integrity and contingency similarity, discuss existing challenges, and identify potential improvements.eng
dc.language.isoeng-
dc.rights.urihttps://creativecommons.org/licenses/by/4.0/-
dc.subject.ddc610-
dc.titleImproved techniques for training tabular GANs using Cramer's V statisticseng
dc.typeArticle-
local.versionTypepublishedVersion-
local.bibliographicCitation.journaltitleProceedings of the 36th Canadian Conference on Artificial Intelligence-
local.bibliographicCitation.pagestart1-
local.bibliographicCitation.pageend12-
local.bibliographicCitation.publishernameCanadian Artificial Intelligence Association-
local.bibliographicCitation.publisherplace[Kitchener, ON]-
local.bibliographicCitation.doi10.21428/594757db.4c0ffb71-
local.openaccesstrue-
dc.identifier.ppn1870547667-
cbs.publication.displayform2023-
local.bibliographicCitation.year2023-
cbs.sru.importDate2024-03-04T08:47:03Z-
local.bibliographicCitationEnthalten in Proceedings of the 36th Canadian Conference on Artificial Intelligence - [Kitchener, ON] : Canadian Artificial Intelligence Association, 2023-
local.accessrights.dnbfree-
Appears in Collections:Open Access Publikationen der MLU

Files in This Item:
File Description SizeFormat 
51682622311607.pdf844.76 kBAdobe PDFThumbnail
View/Open