Aller au contenu principal
Accès ouvert déclaré2026article

Evaluating Baseline Architectures for English-magahi Machine Translation: Architectural Tuning and the Limits of Generalization on Small Parallel Corpus

0Citations signalées
1Institutions associées
1Pays d’affiliation

Résumé fourni par la source

NMT models for extremely low-resource Indic languages suffer from severe data bottleneck problems. In this paper, we analyze the potential and limitations of architectural fine-tuning approaches for English-Magahi translation under the strict data limit of a parallel corpus of 13840 data pairs. We developed two models from scratch: a simple word-level CNN-BiLSTM and an improved subword BiLSTM with BPE tokenization, label smoothing, scheduled sampling, and attention masking. Despite employing sophisticated regularization techniques, our models exhibit a bizarre problem of training failure. Both our models produce near perfect performance (>90 BLEU) on template-only held-out test set but fail terribly (<1 BLEU) on out-of-distribution examples. We identify the reason for this generalization failure as "template collapse" - a situation wherein the model learns artificial target structures rather than cross-lingual mappings. This work shows the empirical upper bound on the performance of NMT models in data-sparse situations. Neither architectural adjustments nor tokenization changes can replace the need for larger and diverse datasets. Future works on robust Magahi NMT models will have to go beyond from-scratch initialization and adopt multi-lingual transfer learning and synthetic data augmentations.

Institutions

Sujets associés

Natural Language Processing TechniquesTopic ModelingBig Data and Digital Economy

BNTIC News n’est pas le producteur de ces données. Métadonnées interrogées à la demande auprès de OpenAlex (CC0). Sources et limites.