Aller au contenu principal
Accès ouvert déclaré 2026 article

Identifying and Protecting Records at Risk for Membership Inference Attacks Against Synthetic Data

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

With increasing interest in leveraging sensitive data, such as health or finanical, for machine learning and other data analysis tasks, the demand for privacy-preserving techniques increases as well. Synthetic data, which is artificially generated by a model trained on real data, as such a technique for micro-data is gaining popularity, especially due to its ability to maintain data utility while reducing disclosure risks, as the observations in the synthetic data do not directly correspond to any individual in the original dataset, making it less susceptible to record linkage or re-identification. Despite this advantage, recent studies have demonstrated potential risks related to membership disclosure, which can occur through attacks that try to determine if a specific record was used to train a model when publishing synthetic micro-level data. This paper explores the potential of synthetic data as a solution to privacy-preserving data publishing by quantifying the risk of membership disclosure; we specifically assess whether outliers are more vulnerable to this type of attack. Furthermore, we analyse whether removing records that are at high risk for membership inference attacks from the training set is effective as a mitigation, and quantify the utility-privacy trade-off introduced. We evaluate foru different synthesizer methods, including Bayesian Networks, Gaussian Copula, CTGAN, and TVAE, on four different datasets. We find that the risk of membership inference varies considerably across synthesizers and datasets, and that the commonly held assumption that outliers are universally more vulnerable is not supported by our evaluation, particularly for distance-based attack approach.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

La source scientifique ouverte est momentanément indisponible.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Privacy-Preserving Technologies in DataData Quality and ManagementImbalanced Data Classification Techniques

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref et Europe PMC, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune réponse conservée. Sources et limites.