Aller au contenu principal
2026 article

A Zero-Training Data Cleaning System With Large Language Models

0Citations signalées — pas une note de qualité
2Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

Data cleaning (DC) is a crucial yet challenging step for many data engineering tasks. Traditional pre-configuration DC methods rely heavily on predefined rules or constraints, demanding significant domain knowledge and manual effort. While configuration-free DC approaches have been explored, they still demand extensive feature engineering or labeled data for intensive model training. In this paper, we propose azero-training and interpretable DCsystem, named${\sf ZeroDC}$, that leverageslarge language models(LLMs) to generate data cleaning rules and chain-of-thoughts (CoTs), without the need for model training.${\sf ZeroDC}$consists of two modules,iterative detection rule generation(IDG) andtraining-free explainable correction(TEC). To generate high-quality error detection rules with minimal human feedback, IDG first bootstraps a set of rules viacontrastive rule initiationon sampled syntactic and semantic contrastive pairs. It then progressively enhances them through aniterative rule refinementworkflow that selects the most informative elements for updates. TEC constructs acontextual-relevant tuple retrieverusing aweighted cosine similarityfunction to efficiently identify the most relevant tuples for each dirty value, reducing redundancy in the LLM prompts and lowering computational costs. It further prompts for generatingcorrection CoTsfor user-corrected representative values, as well as prompts for creatingcorrection rulesandexplainable corrections, which automatically provide explanations for correction results, all without the need for model training. Extensive experiments conducted on various real-world datasets demonstrate that${\sf ZeroDC}$achieves, on average, a 5.36% increase in accuracy and an 8.16x speedup compared to state-of-the-art methods. The codes and datasets of this paper are available athttps://github.com/YangChen32768/ZeroDC.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
A Zero-Training Data Cleaning System With Large Language Models
Date Crossref
01/04/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Data Quality and ManagementResearch Data Management PracticesPrivacy-Preserving Technologies in Data

BNTIC News n’est pas le producteur de ces données. Recherche à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, ROR et la Banque mondiale, sans clé ; OpenAlex reste optionnel. Aucun service payant requis, aucune donnée externe enregistrée en base. Sources et limites.