Aller au contenu principal
Accès ouvert déclaré 2026 article

Recommendation: How to Use Synthetic Data in Machine Learning or Decision Support

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : gb. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Data serves as the foundation of contemporary artificial intelligence (AI) systems, yet ethical and practical constraints often limit the availability and usability of real-world datasets. Synthetic data (SD) has emerged as a valuable solution, enabling the development and training of AI models without compromising privacy standards or ethical guidelines. However, many challenges remain to address, from generating high-quality SD using various approaches to investigating the impacts of data training on machine learning (ML) models. This study examines the impact of balancing real and SD and provides some recommendations that researchers can further utilise to improve the ML model’s training process. Three datasets, Mobile Health (MHealth), High-Energy Physics Mass (HEPMass), and US Company Bankruptcy Prediction (UCBP), were pre-processed to ensure compatibility and used as the basis for SD generation using Conditional Tabular Generative Adversarial Networks (CTGAN) and Tabular Variational Autoencoders (TVAE). The study employed a hybrid data generation approach, splitting the training data into varying proportions of real and SD, with performance evaluated through 5-fold cross-validation on Decision Tree (DT), Gaussian Naive Bayes (GNB), and Linear Support Vector (L-SVM) ML models. The results indicate the optimal balance of real and SD for maximising model performance. Analysis of benchmarking results across three datasets shows that combining 30% real data with 70% CTGAN-generated synthetic data achieves the highest accuracy and overall model performance. In contrast, when using TVAE-generated data, a 20% real and 80% synthetic split is recommended to maintain similar performance. Additional experiments conducted on 10% reduced subsets showed that while the primary trends persisted under limited-data conditions, the optimal real-to-synthetic data ratio became more sensitive to the specific dataset. Statistical significance of the model performance differences was further confirmed using paired t-tests across all evaluated mixing ratios. This research analyses these results and considers the state of the art to recommend how further synthetic data can be useful with machine learning models and which approaches to adopt.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Recommendation: How to Use Synthetic Data in Machine Learning or Decision Support
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Explainable Artificial Intelligence (XAI)Recommender Systems and TechniquesMachine Learning and Data Classification

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.