AnchorVLN: Geometry-Anchored Vision-Language Grounding for Open-Vocabulary Navigation

IROS 2026 · Bridging Perspectives in Navigation · 2nd Place Real-Robot CMU VLN Challenge
Long Giang Vu* Chengkai Yao* Yuxin Liu Rajath C. Aralikatti FNU Aryan

* Equal Contribution

① A semantic–geometric decoupling enforced by the MCP schema - no tool takes a metric argument, so the VLM does language-to-pixel grounding and LiDAR produces the metres.
② Open-vocabulary perception at high frequency (scene graph at 2 Hz) paired with low-frequency VLM reasoning, enabling the agent to detect and ground novel objects without a fixed vocabulary.
③ Decision-level reasoning - the VLM composes tool calls at inference time, replacing hand-scripted per-tick planners.
④ A unified, portable framework: the same MCP interface runs on different robot platforms without rewriting the control stack.

Abstract

We address Vision-Language Navigation in unseen indoor environments where robots must follow natural-language instructions and locate referred objects without pre-built maps or fixed vocabularies. VLMs excel at semantic understanding but cannot reliably emit metric quantities such as range and bearing. We implement this separation as Embodied-Nav-MCP, a Model Context Protocol server whose 10 callable tools inherently enforce metric boundaries -no tool accepts distance in meters or bearing in radians -preventing hallucinated coordinates from reaching the robot. On the CMU VLN Challenge 2026 development set, AnchorVLN achieves 64.4% instruction-following accuracy and clears the object-reference overlap threshold on 10 / 45 questions versus 0 for direct estimation.

Key Results

64.4%
Instruction Following
10/45
Object Reference Hits
2.48m
Median Center Error
2nd
Real-Robot CMU VLN Challenge

System Architecture

AnchorVLN system architecture: Scene graph builder feeds into NAV SYS with VLM client (reasoning) separated from Embodied Navigation MCP (geometry) by a dashed boundary

Figure 1. AnchorVLN architecture. The VLM client reasons about what to look for (phrases); the Embodied Navigation MCP decides where it is (geometry). No metric value crosses the dashed boundary - only phrases and opaque handles.

The MCP Interface

Every input is a phrase or a handle from an earlier call - no input is a number - so the client can name a measurement the system has made but cannot author one.

ToolInputReturns
Perception - query or steer; robot stays put
groundphrasehandles, or refusal
expand_vocabphraseclasses added
verify_presentphrasehandle, or not found
nearesthandles, anchorhandle
farthesthandles, anchorhandle
betweentwo handleshandle (a passage)
ToolInputReturns
Motion - the vehicle moves
drive_tohandlearrived, moved, stalled
explorehandles to avoidhandles newly seen
orbithandlesame handle, refined
Output
publishhandle -

Experiment Setup

Three of the fifteen CMU VLN Challenge scenes

CMU VLN Challenge 2026 - 15 indoor scenes, 10-min budget per question.
Instruction following: drive a free-form language instruction; scored on the trajectory. Object reference: find the referred object and answer with its 3D bounding box.

3D scene graph of livingroom_1  - dozens of objects as coloured point clusters in map coordinates

Each colour is one object record - id, label, 3D box - fused from LiDAR sweeps at 2 Hz and refined as the robot moves. Every handle the VLM client uses points into this graph.

Results

ArmScoreΔtnDTWTLApproach
Full system64.4%––0.3601.231.58 m
w/o Lift63.6%−0.80.230.2791.531.73 m
w/o Platform Model51.1%−13.32.770.3520.792.10 m
ArmHit (IoU ≥ 0.1)Mean IoU
Full system (grounded)22%0.060
w/o tool structure (naive)18%0.054
w/o geometric grounding0%0.000

MCP Tool Illustration

Recorded steps traced through MCP tools: the VLM reasons about what; perception measures where.

arabic_room "Drive to the potted plant furthest from the hookah" MEASURED
Back view: potted plant by doorway, boxed in red
back · target
Front view: second plant and hookah on table
front · anchor
1
ground's VLM names 2 candidate plants and the hookah anchor.
2
Each box is lifted through the LiDAR scan into a handle.
3
farthest(plants, hookah) → the plant by the doorway.
4
drive_to → settles, moved 1.10 m.
VLM determines candidate objects. MCP leverages perception toolbox and compares candidate distances.
office_2 "Drive to the window closest to the clock" REFUSED · IMPLIED SIZE
Front view: candidate windows in orange, chosen in red
front · candidates
Right view: wall clock boxed in blue
right · anchor
1
ground's VLM names candidate windows and the clock anchor.
2
nearest(windows, clock) → the right-hand sash window.
3
Implied-size gate: 28.15 m through glass ⇒ a 15.88 m window. Lift refused.
4
drive_to takes a capped step (0.93 m) toward the window, then looks again.
LiDAR passes through the glass and realizes range is too far. The box implies a 15.88 m window. Refused.

Demo Videos

Across all examples, ground resolves objects the scene graph has not yet detected, and drive_to is the fundamental motion primitive.

Living room: Go to the dining table, then go to the fireplace.
The VLM decides at each step whether the current viewpoint is sufficient to commit to an object or whether further exploration is needed. When it realizes the camera elevation is too low, it autonomously changes viewpoint before continuing.
Living room: Find the horse figurine above the fireplace.
The agent uses explore to collect all candidate objects across the scene, then selects the correct one using the anchor object (the fireplace) via nearest.
Living room: Find the fossil decoration on the bookcase.
The initial vocabulary does not include "fossil." The VLM calls expand_vocab to add it to the detector, enabling the scene graph to localize the object on the next perception tick.
Japanese room: Find the vase closest to the zen stone decoration.
Combines expand_vocab for "zen stone" with orbit around the candidate to refine the 3D bounding box before committing to a localization answer.

Ablation Studies

"the table closest to the dish"

Full system with explore HIT · IoU 0.32
Full system: picks the correct low table beside the dish Full system scene graph: many objects, two table candidates
Naive no explore MISS · IoU 0.00
Naive: picks the only table in its sparse graph -the wrong one Naive scene graph: few objects, single table candidate

"the fossil decoration closest to the big sofa"

Full system with expand_vocab HIT · IoU 0.32
Full system: grounded as nautilus shell sculpture
Naive fixed 100-class prior MISS · IoU 0.00
Naive: falls back to ammonite artwork, picks wrong instance

Cross-Platform Portability

🤖
The same Embodied-Nav-MCP interface runs on a humanoid robot - only the controller model is replaced. The VLM client, tool schemas, and scene graph are unchanged.
Apartment: Go to the dining table, then go to the refrigerator.
Basic sequential navigation using ground and drive_to - the same tool calls that work on the wheeled platform transfer directly to a humanoid.
Apartment: Find the kettle on the kitchen counter.
The agent uses orbit around the kitchen counter to refine the kettle's 3D bounding box before committing to the localization answer.

BibTeX

@article{vu2026anchorvln,
  title   = {AnchorVLN: Geometry-Anchored Vision-Language
             Grounding for Open-Vocabulary Navigation},
  author  = {Vu, Long Giang and Yao, Chengkai and Liu, Yuxin
             and Aralikatti, Rajath Chandrashekar and Aryan, FNU},
  journal = {arXiv preprint arXiv:2609.12285},
  year    = {2026}
}