TASK 26 — Can a VLM ground a referring expression, measured¶
TASK 23–25 established the ceiling our own perception stack runs into: the fused object map holds the objects the fourteen instruction-following questions name only 67 % of the time, and 74 % of the misses are words absent from the 110-class list rather than detection failures. Instruction following is ~70 % of the challenge score and is answered with waypoints, so that ceiling caps the thing worth most.
docs/vlm_grounding_test.md proposed the experiment that decides whether to
keep that architecture. This is the result.
Headline: a VLM looking at a single frame from the start pose localises 70–72 % of named objects, beating the 67 % our full-coverage tour reaches, with a median error of 0.11 m against a 0.58 m budget. After two geometric gates the false-positive rate lands at 11.1 %, against a 10 % acceptance bar.
What was run¶
| Test set | the same 60 (question, named object) pairs scripts/retrieval_probe.py scores, deduplicated to 54 unique (scene, phrase) |
| Frame | the first frame of each recorded corpus — the start pose, no driving |
| Input | the four 640×640 perspective faces the sidecar already unwraps for YOLO |
| Model | claude-opus-5 (the Gemini key does not exist yet; Claude's boxes proved accurate enough that Part B did not need it) |
| Hit criterion | lifted position within 1.0 m of any ground-truth instance of the phrase's synonym family — identical to retrieval_probe |
Pipeline per pair: referring expression → VLM → box_2d → face pixel → camera
ray → lidar returns in an adaptive cone → dominant depth cluster → map-frame
position → error against ground truth.
Results¶
| VLM, start pose | object map, start pose | object map, full tour | |
|---|---|---|---|
| localised within 1.0 m | 72.2 % | 48.3 % | 67 % (k=3) |
| median error of hits | 0.11 m | 0.236 m | — |
| p90 error | 0.52 m | — | — |
| false positives | 24.1 % → 11.1 % gated | — | — |
The error budget is 0.58 m, the median distance the reference trajectories keep
from the objects they name (scripts/traj_tolerance.py). p90 sits just inside
it.
Two prompt bugs that would have inverted the conclusion¶
Coordinate convention. The first run parsed the boxes as the 0–1000
normalised values the prompt asked for and measured a 23.7° bearing error, with
the box drawn on a wall. The model had answered in raw pixels. On a
640-pixel face both readings produce numbers in the same range, so nothing in
the data distinguishes them — the correct reading gives a 0.03° bearing
error. Fixed by requiring the model to declare coord_space explicitly. Had
this gone unnoticed the whole approach would have been recorded as a failure.
A prompt written for only half the test set. v1 demanded a "distinguishing
feature" before reporting a match, which is right for the tea table with the
elephant figurine on it and wrong for chair, table, tv — bare nouns with
no qualifier, which the model then dutifully refused. Splitting the instruction
into a bare-noun branch and a compound branch moved chinese_room from 33 % to
56 %.
What the false positives actually are¶
scripts/vlm_fp_report.py renders every failure with the model's box in yellow
and ground truth projected into the same face in cyan
(artifacts/fp/false_positives.jpg). Reading the panels rather than the error
column changes the diagnosis:
| cause | n | what it looks like |
|---|---|---|
| no lidar return | 4 | ground truth inside the box; the cone is empty even at 14° |
| ray overshot | 3 | open stair, pendant lamp — beam passes through and finds the far wall |
| ray stopped short | 2 | box correct, beam measures an occluder in front |
| wrong instance | 2 | balcony chairs through glass; a side table read as the tray's table |
| truth not in this face | 2 | the elephant figurine answered for the horse; a TV stand for the cabinet |
Nine of the thirteen false positives are cases where the model identified the object correctly and the depth lift failed. Counting only genuine misidentifications gives 4/54 = 7.4 %, inside the acceptance bar.
A sampling artefact was suspected inside this too: one registered scan carries
~976 points per steradian, so a 2° cone expects under four returns and
routinely catches none. locate() now widens the cone until it has enough to
take a median of, and reports which cone it used. That fixed the tea table but
not the four empty cones above, whose real cause TASK 27 traces to a
forward-tilted lidar with a bearing-dependent blind cone.
Three proposed fixes for the depth problem, and what measurement said¶
| proposal | verdict |
|---|---|
| (1) reject lifts beyond a distance cutoff | refuted. False positives are closer than hits (median 2.52 m vs 2.92 m); 7 of 9 sit under 3 m. No threshold separates the distributions — a 3 m cutoff discards 45 % of hits to remove a third of the errors. |
| (2) let the VLM judge whether the scanner can be trusted | refuted. Asked as visual facts rather than sensor questions, same_space returned true for all 52 claimed sightings — including the chairs on the far side of a glass door. occlusion pointed the right way (clear 77 %, see_through 62 %) but on n=8 with no separable margin. |
| (3) let the VLM estimate distance by perspective | partly true, and redundant. Raw estimates are biased 1.81× long with only 21 % within 1.5×. The bias is systematic: dividing it out lifts that to 65 % (in-sample — the factor was fit on this data and needs a held-out check). Even debiased it rejected nothing the geometric gates had not already caught. |
Recommending (2) as "most feasible" was wrong, and the data said so within one run. In hindsight the reason is structural: all three failure modes — glass, open structure, occluder — surface as one contradiction, between the measured range and the size the object would have to be at that range. That contradiction is a multiplication we can do ourselves, and asking a model to notice it introduces judgment where arithmetic suffices.
The gates that worked¶
| gate | hits | false positives | cost |
|---|---|---|---|
| none | 72.2 % | 24.1 % | — |
| A — decline when the cone is empty | 72.2 % | 16.7 % | zero |
| B — implied size within [0.15, 4.0] m | 70.4 % | 18.5 % | 1 hit |
| A + B | 70.4 % | 11.1 % | 1 hit |
| A + B + debiased VLM distance | 70.4 % | 11.1 % | 1 hit |
implied_size = 2 · lift_range · tan(box_angle / 2). Gate A is free: no
threshold, no model, and refusing to answer without depth data is strictly
better than answering from none.
Gate B uses a crude global band. Replacing it with the per-class priors in
perception/size_prior.py looked like the obvious next step; TASK 27 tried it
and it loses at every tolerance — 12 of 52 phrases have no prior, and every
surviving false positive is a wrong instance whose size matches the right one.
Bearing on the architecture¶
The pseudocode reviewed this session survives, with three corrections that came out of the measurements:
- Two loops, not one. A VLM call costs 1–3 s while the reference paths are ~10 m of driving. Geometry runs at control rate; the VLM fires on events — start, arrival, end of an exploration leg. Five to ten calls per question.
boundis notdetected. A detection is per-frame; a binding is a committed map entry that survives out of view. Everything downstream reasons over the binding table, never over the last reply.- Geometry validates. The gates above are that step, and without them the claim rate is 96 % against a 72 % hit rate.
Separately, avoid turns out to cover 3 of 30 instruction questions while
10 use take the path between …, a required passage. Treating every
between as a keep-out would lose ten questions to protect three.
scripts/keepout_radius.py also shows the reference path drives straight
through four of the six chair↔screen corridors in chinese_room q5, so "the path
between A and B" can only mean the closest pair — a geometric
disambiguation, not a semantic one.
Cost¶
Measured, not estimated: 2 990 input tokens (four faces ≈ 2 400) and 460 output tokens per call. At Opus 5's $5/$25 per MTok that is $0.0265 a call — $1.43 for a 54-pair sweep, and roughly $4 for a full competition run at ten calls across fifteen questions. Cost is not a constraint here; the 1–3 s latency is.
Artefacts¶
| path | what |
|---|---|
scripts/snap.sh |
capture image + scan + pose from the running sim |
scripts/grab_faces.py |
equirect → the sidecar's four faces, via its own LUTs |
scripts/vlm_probe.py |
one grounding call; the prompt lives here |
scripts/vlm_locate.py |
box → ray → lidar → metres, scored against a GT centre |
scripts/vlm_sweep.py |
the 54-pair sweep; replies cached, so re-runs are free |
scripts/vlm_fp_report.py |
annotated panel per false positive |
scripts/vlm_gates.py |
offline gate sweep — no API calls |
scripts/keepout_radius.py |
two-sided bounds on the avoidance radius |
artifacts/vlm_sweep.json |
per-pair rows |
artifacts/fp/false_positives.jpg |
the twelve failures, drawn |
Open¶
- ~~Gate B on per-class size priors rather than a global band.~~ Refuted in TASK 27.
- Held-out validation of the 1.81× distance bias, if (3) is revisited.
- Re-run without the (scene, phrase) deduplication, for a strictly identical
denominator to
retrieval_probe's 60. - Everything here is measured from the start pose on one frame. Distance, occlusion and out-of-vocabulary hard cases are untested; so is Gemini, whose detection output is a trained capability rather than an emergent one.