A Comprehensive Survey on Data Distillation: Techniques, Frameworks, and Future Directions
Rattachement africain : in. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
The increased adoption of machine learning techniques has led to exponential growth in data generation and utilization. This growth has necessitated efficient storage, processing, and utilization of this data, which presents critical challenges, particularly in resource-constrained environments such as the Internet of Things (IoT) and edge devices. Data distillation has emerged as a promising solution that reduces dataset size while preserving essential information and optimizing computational resources. This survey provides a comprehensive analysis of data reduction techniques, covering methodologies such as knowledge distillation, coreset selection, hyperparameter optimization, and generative modeling. We further explore various data distillation learning frameworks, including performance, gradient, parameter, and distribution matching, highlighting their effectiveness in different data modalities such as images, graphs, and text. Furthermore, we examine the implications of data distillation in key areas such as continual and federated learning, privacy preservation, security, healthcare, IoT applications, and edge computing. By enabling lightweight models with minimal computational overhead, data distillation facilitates real-time inference and decision-making on edge devices, making it highly relevant for low-power, bandwidth-limited environments. Data distillation offers numerous advantages in improving model efficiency, reducing training costs, and enhancing privacy. However, data distillation faces numerous challenges related to scalability, computational complexity, and information retention. This survey identifies these challenges and outlines potential future research directions, providing insights for researchers seeking to leverage data distillation for scalable and efficient machine-learning applications.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Comprehensive Survey on Data Distillation: Techniques, Frameworks, and Future Directions
- Date Crossref
- 01/02/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.