Aller au contenu principal
Accès ouvert déclaré 2026 article

Speaker-Mediated Route Knowledge Transfer for Goal-Oriented Vision-and-Language Navigation

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Goal-oriented vision-and-language navigation (VLN) requires an agent to reach a target from a high-level instruction that does not prescribe a route. Training therefore faces two problems. First, terminal and distance-based feedback leave partial trajectories without route-semantic supervision. Second, a single annotated path introduces single-reference route bias because several routes may satisfy the same goal. We propose Speaker-Mediated Route Knowledge Transfer (SMRKT), which comprises three components. To provide the route knowledge required by both problems, we trained a trajectory-to-language speaker on route-following VLN data and froze it as a training-time evaluator in the cross-granularity route-prior transfer process. To address the lack of route-semantic supervision, Speaker Reconstruction Progress Reward converts consecutive reconstruction-loss changes into intermediate feedback. To address single-reference route bias, Disagreement-Aware Trajectory Reuse interprets the speaker score with terminal outcome, retaining successful low-score paths as alternative demonstrations and failed high-score paths as hard replay cases. The navigator retains the original goal-oriented instruction, and the speaker is absent at inference. In DUET-based validation on two goal-oriented VLN datasets, SMRKT improves Success Rate (SR) by 1.42 percentage points on the unseen validation setting of the REVERIE dataset, by 1.32 percentage points on the Unseen Houses split of the SOON dataset, and by 2.22 percentage points on the Unseen Instructions split of the SOON dataset; SPL and OSR rise in some settings and fall in others. The results support transferred route knowledge as a training signal, with environmental success retained as the final task criterion.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
Speaker-Mediated Route Knowledge Transfer for Goal-Oriented Vision-and-Language Navigation
Date Crossref
01/09/2026
Éditeur
MDPI AG
Type
journal-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Beijing University of Technology Beijing Key Laboratory of Multimodal Cognitive Computing and Intelligent Software Technology pays non établi dans la notice
    Université ou école supérieure

Beijing Key Laboratory of Multimodal Cognitive Computing and Intelligent Software Technology — Beijing University of Technology.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsSpeech and dialogue systemsNatural Language Processing Techniques

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.