TASK 39 — Question-directed VLM navigation to the referenced object (Opus 5)¶
Purpose¶
A second VLM navigation mode, for the first question type
(object_reference, e.g. "Find the red chair" / "the pillow closest to the
sushi"). Where TASK 38's nav_vlm explores undirected, this mode is
goal-directed: given the question, drive the robot to the object it names,
then answer the question from the built scene graph.
The three constraints from the request map onto the pipeline as follows:
- Scene graph still built at 2 Hz — the app tick loop already calls
responder.ingest()every tick during the exploration phase; the newscene_clauderesponder runs the perception detect → lift → fuse cycle there, unchanged. - Claude called only on waypoint reached or skipped — the app's existing
reach / skip / stuck-watchdog supervisor already drives
explorer.advance()/force_skip()on exactly those events. The new explorer's model call fires on those hooks (inherited from TASK 38's event-driven trigger), never on a timer. - At arrival, dump the scene graph to Claude to answer — when the model
declares arrival the explorer completes, the tick loop hands over to
scene_claude, which serialises the full scene graph and asks Claude to pick the answerobject_id.
What was built¶
Reuses TASK 38's nav_vlm engine/config/render and the app's navigation
supervisor — the new surface is small.
| Piece | Role |
|---|---|
nav_vlm/engine.py |
Generalised AnthropicNavEngine: a call_tool(system, tool, user_text, images) primitive with injectable system_prompt/tool; propose() now sits on top of it. Backward-compatible (TASK 38 tests unchanged). |
nav_vlm/task1_prompts.py |
The goal-directed nav prompt + propose_or_arrive tool (done = arrived at the target object), the object-reference answer prompt + answer_object_reference tool, and the user-text / scene-summary builders. |
exploration/_nav_task1.py |
NavTask1Explorer(NavVLMExplorer) — holds until a question arrives, feeds the question + objects-detected-so-far into each call, and completes when the model declares arrival. Inherits grid, reachability snapping, async off-thread calls, and the reach/skip/watchdog machinery. |
scene_claude/responder.py |
SceneClaudeResponder — builds the scene at 2 Hz via a wrapped PerceptionResponder (ingest), and at arrival dumps scene.to_dict() + a panorama + occupancy map to Claude to select the answer object_id → ObjectReferenceResponse. Non-object_reference questions fall back to the perception responder. |
app/main.py |
Wires XIAO_HEI_EXPLORATION_STRATEGY=nav_task1 (threading the shared scene into _build_explorer) and XIAO_HEI_RESPONDER=scene_claude. |
Design — the loop¶
question arrives ─► NavTask1Explorer (goal-directed)
every tick: PerceptionResponder.ingest → scene graph grows (2 Hz)
on reach (advance) / skip (force_skip / watchdog): ONE Claude call
→ propose_or_arrive: next waypoint toward target, OR done=arrived
arrived ─► explorer.is_complete() ─► SceneClaudeResponder.respond
→ dump scene graph + images to Claude
→ answer_object_reference(object_id) ─► ObjectReferenceResponse
Before the question arrives the explorer holds (no target object to head for) while the scene graph still fills in. Every waypoint the model proposes is snapped to a grid-reachable free cell (TASK 38's anti-wedge safeguard), and a cannot-reach outcome feeds its reason into the next prompt.
If the target object is not yet in the scene graph, the model is told so (the prompt lists the detected objects) and proposes search waypoints toward unobserved area until it appears — then approaches and declares arrival. Arrival is Claude's decision, not a geometric threshold.
Robustness¶
- Answer always emitted. If Claude fails, times out, or names an id not in
the graph, the responder retries a few answer ticks (the scene may still be
settling) and then falls back to the closest object whose label appears in the
question — so the run never ends without an
ObjectReferenceResponse. - Fails fast on a missing key for the answerer (
scene_clauderaises at boot); the explorer fails soft (disables, likenav_vlm). - No SDK / key / GPU needed to build or test — lazy
anthropicimport, injectable fake client/engine.
Tests (no key / no network)¶
tests/test_nav_task1_explorer.py— holds before a question; a question triggers nav and the prompt carries the question + scene graph; model arrival completes; reach re-triggers; skip feeds the failure forward.tests/test_scene_claude_responder.py— a valid model pick becomes theObjectReferenceResponse;ingestbuilds without answering; invalid-id retry → label fallback; engine failure tolerated; non-object_referencedelegates to perception.tests/test_nav_vlm_engine.py— added acall_toolcustom-tool test.
29 mode tests pass; the full suite is 524 passed (the 2 collection errors are pre-existing missing optional deps: networkx/scipy).
To go live¶
- Provide
ANTHROPIC_API_KEY(env or repo.env). uv sync --extra nav-vlm(addsanthropic; matplotlib/pillow already needed by the occupancy render).- Select both halves:
XIAO_HEI_RESPONDER=scene_claudeandXIAO_HEI_EXPLORATION_STRATEGY=nav_task1. The perception sidecar must be up (scene-graph building).
Live validation (2026-08-24)¶
With the Anthropic key in .env (XIAO_HEI_ANTHROPIC_API_KEY), both Claude
paths were smoke-tested against the real claude-opus-5:
- nav (
propose_or_arrive) — returned a goal-directed waypoint with the right reasoning ("no red chair detected yet, so move into unobserved area"); - answer (
answer_object_reference) — correctly selected the matchingobject_idfrom a scene-graph dump; warmup()— succeeds (credentials + model + forced-tool + image).
Two API facts were found and handled in the engine/config:
- Opus 5 deprecates
temperature— the request omits it by default (400 otherwise); still settable via env for other models, sent throughextra_bodyso it works across SDK builds (this environment shipsanthropic1.0.0, whosemessages.create()has notemperaturekeyword). - The vision API rejects sub-minimum images ("Could not process image"), so
warmup()uses a 64×64 solid PNG, not a 1×1.
Still pending: a full sim run (perception sidecar + ROS) driving the loop end-to-end.
Status¶
Scaffold complete, unit-tested against a fake engine, and both Claude paths
live-validated against Opus 5. Full sim run pending. Default
responder/strategy are unchanged (dummy / frontier) — this mode is fully
opt-in.