Aller au contenu principal
2025 conference-paper

End-to-End Model for Vision-Language Navigation Based on Pre-Trained Model

0Citations signalées, ce qui n’est pas une note de qualité
1Institutions déclarées
1Pays d’affiliation déclarés

Rattachement africain : cn. Niveau de preuve : code pays fourni par la source.

Le résumé fourni par la source

Aiming at the problem that embodied intelligent mobile robots have difficulty in accurately navigating in unseen environments according to language instructions, this paper proposes an end-to-end model suitable for visual language navigation. The contextual representation of natural language instructions is extracted by the pre-trained BERT model, and then fed into the model's action predictor, which consists of a global predictor and a local predictor. The global predictor obtains high-level information of the scene through the topological graph to make global action planning. The local predictor only focuses on the visual information around the current position to make local action planning. The predictors are all cross-modal information fused through a transformer-based approach. Finally, the output of both predictors is dynamically fused into actions suitable for the robot. The proposed model is evaluated in the Matterport3D simulator to verify its effectiveness and superiority. Experiment results show that the navigation success rate is improved and mean navigation error is decreased compared with previous work, which can effectively improve the generalization ability of robot navigation in unseen environments.

Ce résumé expose les affirmations des auteurs. BNTIC ne l’interprète pas comme une validation indépendante des résultats.

Le contrôle bibliographique ouvert

DOI retrouvé dans Crossref DOI retrouvé ; titre concordant.

Titre Crossref
End-to-End Model for Vision-Language Navigation Based on Pre-Trained Model
Date Crossref
04/07/2025
Éditeur
IEEE
Type
proceedings-article

Ce recoupement confirme des métadonnées liées au DOI. Il ne confirme ni la méthode ni les conclusions de l’étude, et il ne compte pas comme une seconde source scientifique indépendante.

Où se fait cette recherche

  • Nanjing University of Science and Technology pays non établi dans la notice
    Université ou école supérieure
  • School of Automation pays non établi dans la notice
    Université ou école supérieure

Nanjing University of Science and Technology et School of Automation.

Une affiliation ne permet pas de déduire la nationalité d’un auteur.

Les sujets associés

Multimodal Machine Learning ApplicationsAdvanced Image and Video Retrieval TechniquesSpeech and dialogue systems

BNTIC News n’est pas le producteur de ces données. Les publications sont interrogées à la demande dans Crossref, OpenAIRE, DOAJ, Europe PMC, HAL, DataCite, AfricArXiv, ROR et la Banque mondiale, sans clé d’accès. OpenAlex reste optionnel. Aucun service payant n’est nécessaire et aucune donnée externe n’est enregistrée en base. Consulter les sources et leurs limites.