Aller au contenu principal
2026 article

Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation

0Citations signalées, ce qui n’est pas une note de qualité
3Institutions déclarées
2Pays d’affiliation déclarés

Rattachement africain : cn, gb. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Recent large-scale Vision-and-Language Navigation (VLN) models deliver strong accuracy but remain costly to deployment due to heavy parameters and computation. We tackle efficient VLN in two steps. First, we build a high-performing teacher that makes navigation evidence selection explicit and compressible. Concretely, the teacher introduces a small set of learnable query slots to extract global and local action-sufficient navigable evidence from panoramic observations via a Navigable Query Generator, and progressively grounds these evidence tokens to the instruction with an Instruction–Query Aligner for policy prediction. Second, leveraging this explicit query bottleneck as a distillation interface, we train a compact student by transferring both where to attend and what to do: we distill the teacher’s global/local navigable queries with a navigation-aware token-adaptive objective, and further match action distributions during fine-tuning. Experiments on standard VLN benchmarks demonstrate that our student nearly matches the teacher’s navigation performance while reducing the number of parameters by 93.65% compared to the teacher.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Query, Align, and Distill: Navigation-Aware Cross-Modal Interaction for Efficient Vision-and-Language Navigation
Date Crossref
01/01/2026
Éditeur
Institute of Electrical and Electronics Engineers (IEEE)
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Beijing University of Posts and Telecommunications pays non établi dans la notice
    Université ou école supérieure
  • South China University of Technology pays non établi dans la notice
    Université ou école supérieure
  • Aston University pays non établi dans la notice
    Université ou école supérieure

Beijing University of Posts and Telecommunications, South China University of Technology et Aston University.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsSpeech and dialogue systemsRobotics and Sensor-Based Localization

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.