Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
Rattachement africain : us. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Many-shot in-context learning has recently shown promise as an alternative to finetuning, with the major advantage that the same model can be served for multiple tasks.However, this shifts the computational burden from training-time to inference-time, making deployment of many-shot ICL challenging to justify in-practice.This cost is further increased if a custom demonstration set is retrieved for each inference example.We present Dynamic Block-Sparse Attention, a training-free framework for retrieval-based many-shot in-context learning.By combining carefully designed blocksparse attention and retrieval of cached groups of demonstrations, we achieve comparable perexample latency to finetuning while maintaining on average >95% of the best method's accuracy across strong ICL and finetuning baselines.We hope that this will further enable the deployment of many-shot ICL at scale. 1
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Efficient Many-Shot In-Context Learning with Dynamic Block-Sparse Attention
- Date Crossref
- 01/01/2025
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Carnegie Mellon University Language Technologies Institute pays non établi dans la noticeUniversité ou école supérieure
Language Technologies Institute — Carnegie Mellon University.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.