Bitte benutzen Sie diese Kennung, um auf die Ressource zu verweisen: http://dx.doi.org/10.25673/115107
Titel: Improved techniques for training tabular GANs using Cramer's V statistics
Autor(en): Mendikowski, Melle
Schindler, Benjamin
Schmid, Thomas
Möller, Ralf
Hartwig, Mattis
Erscheinungsdatum: 2023
Art: Artikel
Sprache: Englisch
Zusammenfassung: Considering the growing global demand for machine learning training data, synthetic data generation is a reasonable way to address the versatile challenges in data acquisition. Conditional Tabular Generative Adversarial Network (CTGAN), an extension of the widely used Generative Adversarial Network (GAN), is considered one of the most promising techniques in the field of tabular data generation. Despite numerous successes of CTGAN, a lack of preserving categorical dependencies within the data has been identified. In prior work, the Cramer’s V (CV) as a natural metric for representing the correlation of categorical dependencies was proposed for hyperparameter tuning of CTGAN models. In this paper, we explore two novel strategies to directly integrate CV statistics of data batches within CTGAN training. The first approach is a generator loss term that penalizes differences between the CV statistics of the original and generated data. The second innovation is the extraction of the CV matrix as an additional feature for the critic. By applying our proposed methods to three benchmark datasets, we improve the averaged accuracy of supervised learning models trained on synthesized data by 11 % compared to the legacy CTGAN. We also outline the impact of CV statistics on preserving dependencies between categorical data columns in terms of integrity and contingency similarity, discuss existing challenges, and identify potential improvements.
URI: https://opendata.uni-halle.de//handle/1981185920/117063
http://dx.doi.org/10.25673/115107
Open-Access: Open-Access-Publikation
Nutzungslizenz: (CC BY 4.0) Creative Commons Namensnennung 4.0 International(CC BY 4.0) Creative Commons Namensnennung 4.0 International
Journal Titel: Proceedings of the 36th Canadian Conference on Artificial Intelligence
Verlag: Canadian Artificial Intelligence Association
Verlagsort: [Kitchener, ON]
Originalveröffentlichung: 10.21428/594757db.4c0ffb71
Seitenanfang: 1
Seitenende: 12
Enthalten in den Sammlungen:Open Access Publikationen der MLU

Dateien zu dieser Ressource:
Datei Beschreibung GrößeFormat 
51682622311607.pdf844.76 kBAdobe PDFMiniaturbild
Öffnen/Anzeigen