A Hybrid Multimodal Architecture for Named Entity Recognition on Social Media Using Domain-Specific Language Modelling and Token-Level Gated Visual Fusion
Rattachement africain : pk. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Named entity recognition on social media is a markedly harder problem than its counterpart on formal text, owing to the informal vocabulary, creative abbreviations, and context-dependent entity references that characterise platforms such as Twitter. Although attaching an image to a tweet often carries the disambiguating information needed to resolve ambiguous entity mentions, most existing methods either discard the visual channel entirely or incorporate it in ways that introduce noise rather than signal. We present MSCMT_Hybrid, a hybrid architecture that encodes tweet text with BERTweet - a domainspecific language model pre-trained on 850 million tweets – and extracts global visual semantics with a frozen CLIP ViT$\mathrm{L} / 14$encoder. A learnable cross-modal projection maps the image embedding into the text representation space, after which a tokenlevel sigmoid gate selectively modulates visual influence at each token position. A Conditional Random Field decoder enforces globally valid BIO label sequences. We evaluate the model on a merged corpus combining TWITTER-2015 and TWITTER-2017, totalling approximately 12,090 tweet-image pairs. A systematic grid search over learning rate and batch size yields a final weighted F 1 of 75.97 % - a competitive result relative to published state-of-the-art methods on this task, achieved with a substantially simpler cross-modal fusion mechanism than prior Transformer-based approaches. We additionally conduct a crossdataset generalisation analysis and ablation study that identify the individual contribution of each architectural component, and document a systematic failure mode of unified vision-language models when adapted to token-level sequence labelling, providing an empirical motivation for the hybrid design.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- A Hybrid Multimodal Architecture for Named Entity Recognition on Social Media Using Domain-Specific Language Modelling and Token-Level Gated Visual Fusion
- Date Crossref
- 28/04/2026
- Éditeur
- IEEE
- Type
- proceedings-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
National University of Sciences and Technology pays non établi dans la noticeUniversité ou école supérieure
National University of Sciences and Technology.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.