Aller au contenu principal
Accès ouvert déclaré 2025 conference-paper

Knowledge-enhanced Vision-Language Models for Few-Shot Object Detection in Construction Site

0Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : hk, cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Visual understanding of complex construction site objects is critical for project safety management and worker-robot collaboration within the construction domain. However, deploying deep learning algorithms on construction sites presents significant challenges due to high data annotation costs, substantial computational requirements, and the absence of large-scale training datasets. While large-scale pre-trained multimodal foundation models have shown success in natural language understanding and visual recognition, their application in construction safety management remains limited because of the need for domain-specific knowledge. To address these challenges, this paper proposes a knowledge-enhanced multimodal learning approach for few-shot object detection in construction scenarios. The proposed method comprises two components: (1) leveraging existing semantic knowledge in the construction domain to detect potential objects in construction scenes using a template matching approach; and (2) introducing a multimodal image semantic recognition method that integrates visual and textual knowledge specific to the construction field. We evaluate our approach on the AIMDataset. The results demonstrate that, without network training or large-scale construction samples, our method can achieve effective object detection under few-shot conditions using only existing models and a small amount of provided visual and textual knowledge. This approach highlights its potential for applications in construction scenarios.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Knowledge-enhanced Vision-Language Models for Few-Shot Object Detection in Construction Site
Date Crossref
12/08/2025
Éditeur
ABC2 Publishing House
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Occupational Health and Safety ResearchBIM and Construction IntegrationInfrastructure Maintenance and Monitoring

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.