InternRobotics' open platform for building generalized navigation foundation models.
-
Updated
Mar 10, 2026 - Jupyter Notebook
InternRobotics' open platform for building generalized navigation foundation models.
[TPAMI 2024] Official repo of "ETPNav: Evolving Topological Planning for Vision-Language Navigation in Continuous Environments"
[RA-L'25 & ICRA'26] An Reliable and Efficient Framework for Zero-Shot Object Navigation
[CVPR 2024] The code for paper 'Towards Learning a Generalist Model for Embodied Navigation'
[ACM MM 2021 Oral] Official repo of "Neighbor-view Enhanced Model for Vision and Language Navigation"
[AAAI 2026] Official code for "Agent Journey Beyond RGB: Unveiling Hybrid Semantic-Spatial Environmental Representations for Vision-and-Language Navigation"
Official code for VoLN: Vision-Only Long-Horizon Navigation—Paradigm, Benchmark, and Method. An embodied AI UAV benchmark and VoLN-MLLM agent bridging VLN and multimodal LLMs across 7,210 episodes, AirSim, and real-world flights.
[CoRL 2025] Search-TTA: A Multimodal Test-Time Adaptation Framework for Visual Search in the Wild
[Official] [IROS 2024] A goal-oriented planning to lift VLN performance for Closed-Loop Navigation: Simple, Yet Effective
[EACL 2026 Oral] AgentNav: Zero-shot sparsely grounded long-range visual navigation in real-world cities using Multimodal Large Language Models (MLLMs).
Official implementation of Route2Step, which separates semantic progress tracking from local action execution for long-horizon vision-and-language navigation.
TEM Engineering Open Tech!
Streamlit App Combining Vision, Language, and Audio AI Models
Source code and documentation for the ACL 2023 Findings paper "Yes, this Way! Learning to Ground Referring Expressions into Actions with Inter-episodic Feedback from Supportive Teachers"
Re-evaluation of instruction-guided ObjectNav: simple frontier geometry matches or exceeds complex LLM-based navigation pipelines.
Global-Ordering Risk-Aware Vision-Language Navigation for Xiaomi CyberDog2
Offline vision-language navigation for outdoor robots: a 350-instruction decomposition benchmark, evaluations of 21 language models across three GPU platforms, and an open-vocabulary goal detection benchmark.
[T-MM 2026] ReliableNav: Uncertainty-Aware Zero-Shot Vision-Language Navigation in Continuous Environments
Zero-shot Vision-Language Navigation (VLN) pipeline for TurtleBot3 using Gemini 2.5 Flash, YOLO-World, MobileSAM, CLIP and ROS 2 Nav2
To associate your repository with the vision-language-navigation topic, visit your repo's landing page and select "manage topics."