Breaking the Attention Trap in Code LLMs: A Rejection Sampling Approach to Enhance Code Execution Prediction
Résumé fourni par la source
Code-specific Large Language Models (Code LLMs) have greatly improved performance across code-related tasks, offering substantial benefits in practical applications.However, existing research reveals significant performance bottlenecks in Code Execution tasks, which requires models to predict the execution results of given code snippets.This study identifies that the Attention Trap phenomenon in training data constitutes a key constraint on model performance.To address this phenomenon, we propose the Attention Cracking with Rejection Sampling (AC-RS) method.The method first applies structural optimization to training data to eliminate attention traps.Then, it conducts secondary training on the outputs generated by the fine-tuned model to mitigate potential negative impacts from manual data intervention.Experimental results show that AC-RS significantly enhances the accuracy of Code Execution while preserving models' original capabilities.Notably, the optimized 7B model achieves Code Execution accuracy comparable to 32B model and GPT-4o. 1 models with self-generated instructions.In Proceedings of the 61st Annual Meeting of the Association for
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Breaking the Attention Trap in Code LLMs: A Rejection Sampling Approach to Enhance Code Execution Prediction
- Date Crossref
- 01/01/2025
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.