Please use this identifier to cite or link to this item: http://dx.doi.org/10.25673/115107
Title: Improved techniques for training tabular GANs using Cramer's V statistics
Author(s): Mendikowski, Melle
Schindler, Benjamin
Schmid, Thomas
Möller, Ralf
Hartwig, Mattis
Issue Date: 2023
Type: Article
Language: English
Abstract: Considering the growing global demand for machine learning training data, synthetic data generation is a reasonable way to address the versatile challenges in data acquisition. Conditional Tabular Generative Adversarial Network (CTGAN), an extension of the widely used Generative Adversarial Network (GAN), is considered one of the most promising techniques in the field of tabular data generation. Despite numerous successes of CTGAN, a lack of preserving categorical dependencies within the data has been identified. In prior work, the Cramer’s V (CV) as a natural metric for representing the correlation of categorical dependencies was proposed for hyperparameter tuning of CTGAN models. In this paper, we explore two novel strategies to directly integrate CV statistics of data batches within CTGAN training. The first approach is a generator loss term that penalizes differences between the CV statistics of the original and generated data. The second innovation is the extraction of the CV matrix as an additional feature for the critic. By applying our proposed methods to three benchmark datasets, we improve the averaged accuracy of supervised learning models trained on synthesized data by 11 % compared to the legacy CTGAN. We also outline the impact of CV statistics on preserving dependencies between categorical data columns in terms of integrity and contingency similarity, discuss existing challenges, and identify potential improvements.
URI: https://opendata.uni-halle.de//handle/1981185920/117063
http://dx.doi.org/10.25673/115107
Open Access: Open access publication
License: (CC BY 4.0) Creative Commons Attribution 4.0(CC BY 4.0) Creative Commons Attribution 4.0
Journal Title: Proceedings of the 36th Canadian Conference on Artificial Intelligence
Publisher: Canadian Artificial Intelligence Association
Publisher Place: [Kitchener, ON]
Original Publication: 10.21428/594757db.4c0ffb71
Page Start: 1
Page End: 12
Appears in Collections:Open Access Publikationen der MLU

Files in This Item:
File Description SizeFormat 
51682622311607.pdf844.76 kBAdobe PDFThumbnail
View/Open