Speaker-Mediated Route Knowledge Transfer for Goal-Oriented Vision-and-Language Navigation
Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.
Le résumé fourni par la source
Goal-oriented vision-and-language navigation (VLN) requires an agent to reach a target from a high-level instruction that does not prescribe a route. Training therefore faces two problems. First, terminal and distance-based feedback leave partial trajectories without route-semantic supervision. Second, a single annotated path introduces single-reference route bias because several routes may satisfy the same goal. We propose Speaker-Mediated Route Knowledge Transfer (SMRKT), which comprises three components. To provide the route knowledge required by both problems, we trained a trajectory-to-language speaker on route-following VLN data and froze it as a training-time evaluator in the cross-granularity route-prior transfer process. To address the lack of route-semantic supervision, Speaker Reconstruction Progress Reward converts consecutive reconstruction-loss changes into intermediate feedback. To address single-reference route bias, Disagreement-Aware Trajectory Reuse interprets the speaker score with terminal outcome, retaining successful low-score paths as alternative demonstrations and failed high-score paths as hard replay cases. The navigator retains the original goal-oriented instruction, and the speaker is absent at inference. In DUET-based validation on two goal-oriented VLN datasets, SMRKT improves Success Rate (SR) by 1.42 percentage points on the unseen validation setting of the REVERIE dataset, by 1.32 percentage points on the Unseen Houses split of the SOON dataset, and by 2.22 percentage points on the Unseen Instructions split of the SOON dataset; SPL and OSR rise in some settings and fall in others. The results support transferred route knowledge as a training signal, with environmental success retained as the final task criterion.
Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.
Le contrôle bibliographique ouvert
DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.
- Titre Crossref
- Speaker-Mediated Route Knowledge Transfer for Goal-Oriented Vision-and-Language Navigation
- Date Crossref
- 01/09/2026
- Éditeur
- MDPI AG
- Type
- journal-article
Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.
Où se fait cette recherche
-
Beijing University of Technology Beijing Key Laboratory of Multimodal Cognitive Computing and Intelligent Software Technology pays non établi dans la noticeUniversité ou école supérieure
Beijing Key Laboratory of Multimodal Cognitive Computing and Intelligent Software Technology — Beijing University of Technology.
Une affiliation ne permet pas de déduire la nationalité d’un auteur.