Aller au contenu principal
Accès ouvert déclaré 2025 article

Toward Energy-Safe Industrial Monitoring: A Hybrid Language Model Framework for Video Captioning

1Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

In the energy industry, like industrial monitoring scenarios, using generative AI for video captioning technology is crucial in event understanding and safety analysis. Current approaches typically rely on a single language model to decode visual semantics from video frames. Lightweight pre-trained generative models often produce overly generic captions that omit domain-specific details like energy equipment states or procedural steps. Conversely, multimodal large generative AI models can capture fine-grained visual cues but are prone to distraction from complex backgrounds, resulting in hallucinated descriptions that reduce reliability in high-risk energy workflows. To bridge this gap, we propose a collaborative video captioning framework, EnerSafe-Cap (Energy-Safe Video Captioning), which introduces domain-aware prompt engineering to integrate the efficient summarization of lightweight models with the fine-grained analytical capability of large models, enabling multi-level semantic understanding, thereby improving the accuracy and completeness of video content expression. Furthermore, to fully exploit the strengths of both small and large models, we design a dual-path heterogeneous sampling module. The large model receives key frames selected according to inter-frame motion dynamics, while the lightweight model processes densely sampled frames at fixed intervals, thereby capturing complementary spatiotemporal cues global event semantics from salient moments and fine-grained procedural continuity from uniform sampling. Experimental results on commonly used benchmark datasets show that our model outperforms baseline models. Specifically, on the VATEX dataset, our model surpasses the lightweight pre-trained language model SwinBERT by 19.49 in the SentenceBERT metric, and outperforms the multimodal large language model Qwen2-vl-2b by 8.27, validating the effectiveness of the method.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Toward Energy-Safe Industrial Monitoring: A Hybrid Language Model Framework for Video Captioning
Date Crossref
04/12/2025
Éditeur
MDPI AG
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsGenerative Adversarial Networks and Image SynthesisHuman Pose and Action Recognition

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.