A Zero-Training Data Cleaning System With Large Language Models
Résumé fourni par la source
Data cleaning (DC) is a crucial yet challenging step for many data engineering tasks. Traditional pre-configuration DC methods rely heavily on predefined rules or constraints, demanding significant domain knowledge and manual effort. While configuration-free DC approaches have been explored, they still demand extensive feature engineering or labeled data for intensive model training. In this paper, we propose azero-training and interpretable DCsystem, named${\sf ZeroDC}$, that leverageslarge language models(LLMs) to generate data cleaning rules and chain-of-thoughts (CoTs), without the need for model training.${\sf ZeroDC}$consists of two modules,iterative detection rule generation(IDG) andtraining-free explainable correction(TEC). To generate high-quality error detection rules with minimal human feedback, IDG first bootstraps a set of rules viacontrastive rule initiationon sampled syntactic and semantic contrastive pairs. It then progressively enhances them through aniterative rule refinementworkflow that selects the most informative elements for updates. TEC constructs acontextual-relevant tuple retrieverusing aweighted cosine similarityfunction to efficiently identify the most relevant tuples for each dirty value, reducing redundancy in the LLM prompts and lowering computational costs. It further prompts for generatingcorrection CoTsfor user-corrected representative values, as well as prompts for creatingcorrection rulesandexplainable corrections, which automatically provide explanations for correction results, all without the need for model training. Extensive experiments conducted on various real-world datasets demonstrate that${\sf ZeroDC}$achieves, on average, a 5.36% increase in accuracy and an 8.16x speedup compared to state-of-the-art methods. The codes and datasets of this paper are available athttps://github.com/YangChen32768/ZeroDC.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Zero-Training Data Cleaning System With Large Language Models
- Date Crossref
- 01/04/2026
- Éditeur
- Institute of Electrical and Electronics Engineers (IEEE)
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.