Curr-ReFT: Overcoming Training Bottlenecks in Small-scale Vision-Language Models via Curriculum Reinforcement Finetuning
Résumé fourni par la source
State-of-the-art vision-language models (VLMs) require massive scaling that limits practical deployment.Small-scale VLMs offer a practical alternative but face out-of-domain (OOD) collapse when trained with traditional supervised fine-tuning (SFT).Through GeneralPoints experiments, we identify that OOD collapse is due to SFT's tendency to induce visual hallucinations under distribution shifts.Although RL-based post-training effectively mitigates OOD degradation, it faces a critical dilemma with sparse rewards in complex visual reasoning tasks.To this end, we propose Curriculum Reinforcement Finetuning (Curr-ReFT), comprising two sequential stages: (1) Structured Curriculum Reinforcement Learning, which progressively evolves task formats and reward functions to match models' growing capabilities; and (2) Rejected Sampling-based Self-improvement, which maintains the fundamental capabilities of VLMs through selective learning from high-quality examples.Extensive experiments demonstrate that Curr-ReFT achieves state-ofthe-art performance across various visual tasks in both in-and out-of-domain settings and benchmarks.Code and data are available at https://github.com/ding523/Curr_REFT.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Curr-ReFT: Overcoming Training Bottlenecks in Small-scale Vision-Language Models via Curriculum Reinforcement Finetuning
- Date Crossref
- 01/01/2025
- Éditeur
- Association for Computational Linguistics
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.
Institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.