Embodiment in multimodal large language models
Rattachement africain : us, ch. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Multimodal large language models (MLLMs) have demonstrated an extraordinary capacity to bridge textual and visual inputs. Nonetheless, MLLMs still face limitations in situated physical and social interactions in sensorially rich and multimodal real-world settings, where the embodied experience of a living organism appears fundamental. We suggest that the next frontiers for MLLM development require the incorporation of both internal and external embodiment-modeling not only external interactions with the world but also internal states and drives. Here, we describe mechanisms of internal and external embodiment in humans and relate these to current advances in MLLMs in the early stages of aligning to human representations. Our dual-embodied framework proposes to model interactions between these forms of embodiment in MLLMs so as to bridge the gap between multimodal data and world experience.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Embodiment in multimodal large language models
- Date Crossref
- 01/06/2026
- Éditeur
- Elsevier BV
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Les institutions déclarées
Une affiliation ne permet pas de déduire la nationalité d’un auteur.