Aller au contenu principal
Accès ouvert déclaré 2026 preprint

Anthropogenic Regional Adaptation in Multimodal Vision-Language Model

0Citations signalées, ce qui n’est pas une note de qualité
15Institutions déclarées
10Pays d’affiliation déclarés

Rattachement africain : th, ca, my, us, au, ph, gb, sg, id, ae. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

While the field of vision-language (VL) has achieved remarkable success in integrating visual and textual information across multiple languages and domains, there is still no dedicated framework for assessing human-centric alignment in vision-language systems. We offer two contributions to address this gap. First, we introduce Anthropogenic Regional Adaptation: a novel paradigm that aims to optimize model relevance to specific regional contexts while ensuring the retention of global generalization capabilities. Second, we present a simple, but effective adaptation method named Geographical-generalization-made-easy (GG-EZ), which utilizes regional data filtering and model merging. Through comprehensive experiments on 3 VL architectures: large vision-language models, text-to-image diffusion models, and vision-language embedding models, and a case study in Southeast Asia (SEA) regional adaptation, we demonstrate the importance of Anthropogenic Regional Adaptation and the effectiveness of GG-EZ, showing 5-15% gains in cultural relevance metrics across SEA while maintaining over 98% of global performance and even occasionally surpassing it. Our findings establish Anthropogenic Regional Alignment as a foundational paradigm towards applicability of multimodal vision-language models in diverse regions and demonstrate a simple-yet-effective baseline method that optimizes regional value alignment while preserving global generalization.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

Aucun DOI disponible pour le contrôle Crossref.

Où se fait cette recherche

  • Thammasat University pays non établi dans la notice
    Université ou école supérieure
  • Mila - Quebec Artificial Intelligence Institute pays non établi dans la notice
    Structure de recherche
  • Universiti Teknologi Petronas pays non établi dans la notice
    Université ou école supérieure
  • Oracle (United States) pays non établi dans la notice
    Entreprise
  • Monash University pays non établi dans la notice
    Université ou école supérieure
  • John Brown University pays non établi dans la notice
    Université ou école supérieure
  • National University pays non établi dans la notice
    Université ou école supérieure
  • Ateneo de Manila University pays non établi dans la notice
    Université ou école supérieure
  • University College London pays non établi dans la notice
    Université ou école supérieure
  • Nanyang Technological University pays non établi dans la notice
    Université ou école supérieure
  • Universitas Hindu Indonesia pays non établi dans la notice
    Université ou école supérieure
  • University of New Haven pays non établi dans la notice
    Université ou école supérieure

Thammasat University, Mila - Quebec Artificial Intelligence Institute et Universiti Teknologi Petronas, avec 9 autres affiliations.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsLanguage and cultural evolutionCategorization, perception, and language

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.