Aller au contenu principal
Accès ouvert déclaré 2026 conference-abstract

Optimizing Vision-Language Model for Robust Road Damage Assessment via Parameter-Efficient Fine-Tuning

0Citations signalées — pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Résumé fourni par la source

This study investigates the domain adaptation of the Vision-Language Model (VLM) for road damage assessment, focusing on a fine-tuning strategy optimized for resource-constrained engineering environments. Unlike conventional object detection models that operate within fixed label spaces, VLMs provide superior semantic understanding and generalization in complex scenarios. To facilitate practical deployment, this research systematically analyzes key variables of Parameter-Efficient Fine-Tuning (PEFT) to mitigate the high computational demands inherent in large-scale VLMs.In the experimental phase, hyperparameter tuning was conducted using the Low-Rank Adaptation (LoRA) technique. The primary variables included LoRA ranks (16, 32, 64, and 96), training data scale, and image resolutions (1,024ⅹ28ⅹ28 vs. 1,536ⅹ28ⅹ28). A comprehensive dataset of 26,796 images comprising six damage categories and negative samples was established, utilizing a 7n sampling strategy (n=500, 750, 1,000) to address class imbalance. The impact of data volume was evaluated by augmenting the 7,000-sample set (corresponding to n=1,000) to match the full dataset size of 26,796, with zero-shot inference serving as the performance baseline.Experimental results demonstrated substantial improvements over zero-shot inference, indicating that performance positively correlates with increased data scale with augmentation and higher image resolution, while lower LoRA ranks (16, 32) proved most effective for this domain. Furthermore, the introduction of specialized ad-hoc metrics, MmAP and MF1, verified a stable trade-off between False Positives and False Negatives. Notably, to minimize safety-critical False Negatives, a prompt engineering-based 'Double Check' mechanism and multi-turn interactions were utilized. This approach successfully leveraged the model’s inherent reasoning capabilities to refine damage identification through iterative feedback.Acknowledgements This research was supported by Basic Science Research Program through the National Research Foundation of Korea(NRF) funded by the Ministry of Education(RS-2025-25437298)

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Optimizing Vision-Language Model for Robust Road Damage Assessment via Parameter-Efficient Fine-Tuning
Date Crossref
13/03/2026
Éditeur
Copernicus GmbH
Type
posted-content

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude et ne compte pas comme une seconde source scientifique indépendante.

Institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Sujets associés

Infrastructure Maintenance and MonitoringAdvanced Neural Network ApplicationsDomain Adaptation and Few-Shot Learning

BNTIC News n’est pas le producteur de ces données. Exploration à la demande auprès d’OpenAlex, avec contrôle bibliographique public par Crossref. Aucun service payant requis, aucune réponse conservée. Sources et limites.