Enhancing Hallucination Detection in Large Language Models by Internal-External Feature Fusion
Résumé fourni par la source
Detecting non-factual contents in Large Language Model (LLM) generations is the first step for mitigating the hallucination in LLMs. Current probing-based methods rely only on human-annotated labels demonstrating and suffer from limited transferability to out-of-distribution contents, while the consistency-checking based methods fail to correctly detect hallucinations when the LLM consistently produces incorrect outputs. This paper proposes ESIF, which leverages external semantic distribution features to reduce dependency of probing-based methods on data distribution, thereby enhancing the transferability of LLM. In addition, ESIF fuses external and internal features to train the classifier, which improves the accuracy of hallucination detection. These two types of features work in a complementary manner: external features help reduce the reliance on data distribution by capturing the overall semantic trends, while internal features enhance this framework by refining the detection of edge cases in which rely on external features solely may fall short. Experiment results demonstrate that the proposed ESIF outperforms the existing state-of-the-art methods in hallucination detection and achieves excellent transferability.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Enhancing Hallucination Detection in Large Language Models by Internal-External Feature Fusion
- Date Crossref
- 18/10/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.