Aller au contenu principal
Accès ouvert déclaré 2026 article

Development of a multimodal english cross-language translation model based on transformer architecture

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Translation across multiple forms and different languages shows limitations from problems with combining features and problems with moving meaning between languages, and these problems affect how the approach performs in settings that involve various conditions. This study aims to design a unified multimodal Transformer architecture that strengthens cross-linguistic semantic alignment and improves translation stability under heterogeneous input conditions. This study conducted initial processing that combined image and text data, allowing separate encoding of features at higher levels. It then used attention across different forms to combine meaning from multiple sources, producing a representation that showed unity. The approach also introduced space for meaning that multiple languages share at a level beyond surface forms, thereby improving the stability of the mapping between languages in the results. Training without modality and applying noise-enhancement techniques together improved this model’s robustness to incomplete and degraded input conditions that approximate practical deployment scenarios.The model was trained and evaluated on a structured English-German bilingual vision-language dataset composed of aligned image-text pairs collected under controlled experimental conditions.Experimental evaluation in fusion mode yielded BLEU scores of 39.2, METEOR scores of 31.7, and CIDEr scores of 113.8, indicating stronger semantic consistency and improved bilingual generation quality. The results of this research demonstrate that multimodal fusion and cross-language alignment mechanisms offer effective approaches to improve the accuracy and reliability of multimodal translation systems. The findings remain constrained by dataset scale and language coverage, and future research will extend validation across broader linguistic domains and more diverse multimodal scenarios.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Development of a multimodal english cross-language translation model based on transformer architecture
Date Crossref
07/09/2026
Éditeur
Springer Science and Business Media LLC
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Les institutions déclarées

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsGenerative Adversarial Networks and Image SynthesisNatural Language Processing Techniques

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.