Domain-Adapted Captioning Framework for Accessibility and Assistive Technologies
Le résumé fourni par la source
AI-based image captioning has emerged as a cornerstone of next-generation assistive technologies, enabling visually impaired users to access and understand their surroundings through natural language. Despite rapid advances in vision–language models, no existing work has systematically evaluated these systems under the real-world conditions of accessibility data where images are blurred, poorly lit, occluded, and off-centered. This study is the first to benchmark and interpret the performance of modern captioning architectures in open-environment setting. We fine-tune the transformer-based Bootstrapping Language–Image Pre-training (BLIP) model on VizWiz and augment it with object detection and optical character recognition (TrOCR) modules to handle cluttered and text-rich scenes. Caption quality is assessed using BLEU-4, METEOR, and CIDEr metrics, as well as qualitative human-aligned evaluation. Our findings reveal that domain-specific fine-tuning and multimodal integration markedly enhance caption informativeness and contextual accuracy compared with COCO-trained baselines, even under constrained training.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Domain-Adapted Captioning Framework for Accessibility and Assistive Technologies
- Date Crossref
- 17/12/2025
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.