TASK 30 — Four scenes, and the bugs only driving finds¶
TASK 29 built a model of waypointConverter and validated it on one scene. This
is what happened when the loop was rewritten around it and driven on four:
japanese_room, chinese_room, office_1, loft. Roughly fifteen runs.
Every defect below was already in the code before today. None of them showed up in replay. Each needed a robot to move.
Where it ended¶
| scene | target | before | after | notes |
|---|---|---|---|---|
office_1 |
potted plant furthest from the projector screen | 0.89 m | 0.61 m | platform floor is 0.62 m |
japanese_room |
lantern closest to the fan decoration | 1.06 m | 0.99 m | corner; floor is ~1.06 m |
chinese_room |
tea table with the elephant figurine on it | 0.48 m | 0.80 m | seen anywhere in 0.35–0.80 m |
loft |
cup near the TV remote | — | fails | see below |
Distances are vehicle centre to the object's ground-truth box, the same way
traj_tolerance.py measures the reference trajectories, whose median is 0.58 m.
chinese_room is not a regression and not an improvement — it is a spread.
Across the session it landed at 0.35, 0.47, 0.48, 0.52, 0.55 and 0.80 m. The
runs that scored best did so through the blind-cone branch wandering into a good
spot; binding took that branch over and with it the accidental luck. One or two
samples per scene cannot tell a policy change from that spread, and no
conclusion here rests on a single-run comparison.
Seven bugs, in the order driving found them¶
1. The waypoint policy optimised the wrong quantity. It picked the legal
point nearest the target. But the vehicle stops waypointXYRadius short of its
goal along the approach, so of two goals equally far from the target, the one
nearer the vehicle settles further from it. On office_1 two candidates sat
1.088 m and 1.085 m from the target; preferring the first because it was 0.37 m
nearer the vehicle ended the run at 0.89 m instead of 0.69 m. Now scored on
where the vehicle settles. An exact prune (|goal−target| − 0.3 bounds the
best possible settle distance) plus resolving legality once per frame instead of
per query took that from 6 000 ms to 34 ms.
2. The stall test ignored rotation. The platform turns to face its waypoint
before translating — measured: commanded bearing −142.2°, final yaw −146.3°. A
waypoint 138° behind the vehicle therefore produces seconds of pure rotation,
and a position-only stall test reads that as the stack refusing to move:
settled, moved 0.0008 m, while the next frame showed it had turned 56° and was
still turning. This has been present since TASK 28, so some of that task's
"as near as it allows" readings were the loop giving up mid-turn.
3. Arrival still meant "reached our waypoint". That was right when the waypoint was the target minus a standoff. It became wrong the moment the waypoint became "the legal point that settles nearest the target", which can sit metres away: the loop drove to one 2.61 m from the tea table and reported ARRIVED. Reaching a waypoint is now a better vantage point and nothing more.
4. Termination asked whether the converter would move us, not whether moving would get us closer. The legal points around an object form a ring at roughly equal distance from it, so there is always another one worth 1.4 m of driving and 0.08 m of progress. The loop circled the tea table for six calls.
5. …and the fix stopped too early. Judged on gain alone, the first step of
chinese_room declared the platform's floor at 2.15 m, because a 5 m local
terrain map from the start pose has not seen the ground near the object. Driving
anywhere is still worth it there — it is what makes the map grow. Only within
NEAR_M does "cannot improve" mean "cannot get closer". NEAR_M = 1.5 is
calibrated on three scenes, not measured.
6. max_tokens = 2048 truncated the replies. Three runs died on
"unparseable reply", which reads as a model failure and was a budget one: v4 asks
the model to enumerate candidates, anchors and alternates, and a five-lantern
reply overruns 2048 and comes back with the JSON cut mid-object. TASK 26's
measured 460 output tokens was v3. Now 4096, stop_reason == "max_tokens"
raises rather than returning a bad reply, and the raw text is logged on any
parse failure.
7. sim.sh up started a second simulation stack when the requested scene was
already loaded: docker/run correctly does nothing, and the script then ran
system_simulation.sh again inside the same container. Two stacks fought over
/way_point_with_heading — visible only as doubled publisher counts in
preflight — and the drive timed out. up now always tears down first. One run
was discarded.
Target binding, and why the gate is on distance¶
Fixing (3) let the loop reach a second grounding for the first time, which
exposed that it had no memory of what it had committed to. On japanese_room
step 1 bound the lantern 0.19 m from the truth and step 2 produced one 4.18 m
away; the loop drove to it.
The model cannot arbitrate this. It reported same_object_as_previous: True and
higher confidence in both the healthy case and the broken one:
| error, step 1 → 2 | binding moved | |
|---|---|---|
office_1 |
0.56 m → 0.05 m | 0.52 m |
japanese_room |
0.19 m → 4.18 m | 4.37 m |
So the gate is on how far the reading moved, not on identity. That is also what makes it safe against the obvious objection — that binding early locks in an early mistake. It does not block refinement, only teleportation. The failure it accepts is a grossly wrong first sighting, 4/54 = 7.4 % in TASK 26; a jump back is allowed only when the model itself reports a different object at higher confidence, which is untested.
A binding also rescues a lift the blind cone made untrustworthy: the position
was measured once and /state_estimation carries it across the move. On
chinese_room the rejected blind lift read 1.38 m against 1.36 m from binding
plus odometry — 0.02 m apart.
Path constraints: built, not demonstrated¶
ConverterModel(terrain, keepout=[(xy, r)]), with prompt v5 asking for avoid
and gate anchors in the same call as the target.
The geometry is unit-tested and behaves: the allowed set shrinks 836 → 167, a
goal inside a zone is refused, reach_along truncates at the near edge rather
than reading straight through, and — deliberately — snap/settle are
untouched, because they predict the organisers' node and it has never heard of
our constraint.
It has not worked end to end. Across six calls carrying an avoid clause the
model reported an anchor once, and that one lifted to (+4.09, −0.29) — 2.2 m
from the nearest real cabinet, with ceiling as the closest ground-truth object.
n = 1. The radius, 1.2 m, sits inside the 0.86–1.98 m bounds
keepout_radius.py measured.
Refuted in TASK 31. Those bounds came from two of the three keep-out questions; exporting the third inverts them (upper 0.81 m, below the 0.86 m lower bound), and a 1.2 m disc forbids the official reference path in two of the three. The disc is also the wrong shape: the constraint is on the corridor between two anchors, not on the anchors.
gate is parsed and logged, not enforced. It is worth more than avoid: ten of
the thirty instruction questions need a passage driven through, against three
that need a region avoided.
Exploration, and what loft actually is¶
Two changes, neither sufficient.
Reachability. The explore heading was published as a fixed 0.8 m hop without
ever asking whether the vehicle could go that way. On loft the model asked
three times for a heading with 0.00 m of reach while 5.16 m was available 30°
away, and each refusal cost a call. Legs now run as far as the legal set allows,
capped at 3 m, along the drivable direction closest to the one asked for.
This corrects a claim made earlier in the session. Routing exploration through the converter model was blamed for throttling it; measured, the legal set extended 5.16 m and the model's chosen heading was simply blocked.
Memory. Each reply now writes a here clause — its own words for where it is
standing and what it could search — and those are fed back on later calls. In
the model's language, not map coordinates, which it cannot act on.
It did not stop the wandering, because the robot never left a 2 m radius and all six entries describe the same place from different angles.
What loft is¶
The model was right the whole time, and an earlier note in this session claiming it switched from reasoning about the coffee table to the dining table was wrong: it named the dining table in every step of every run. There is a real cup on that table — two, in fact:
| position | on dining table #9 | |
|---|---|---|
cup #107 |
(+1.57, +2.04) | yes |
coffee cup #113 |
(+0.57, +1.56) | yes |
cup #108 — the answer |
(+6.21, +0.68) | no, 5.64 m away |
The loop bound (+0.59, +1.51), which is 0.06 m from #113. The recognition
was accurate. What it could not do was check "near the TV remote", because that
anchor is 0.05 × 0.20 × 0.02 m and returned zero liftable anchors on every call
of every run. So it did the reasonable thing: circled the table trying to get a
clear view of a tabletop that is hidden behind chair backs — and cannot be seen
anyway, from a camera 0.75 m up, standing the 0.75 m away that obstacleDisThre
enforces, at a cup whose base is 0.43 m.
How common is that anchor?¶
scripts/anchor_size.py measures the angular size of every relation's anchor at
the start pose, over the seven scenes with exported ground truth.
Superseded by TASK 31, which repeats this on all fifteen scenes: 92 anchors, 3 % / 10 % / 87 %. The conclusion below holds; the shares shift by a point or two.
| subtends at start | n | share | |
|---|---|---|---|
| 0–2° | 2 | 5 % | tv remote 1.79°, sushi 1.70° |
| 2–5° | 5 | 11 % | resolves on approach |
| > 5° | 37 | 84 % | liftable from the start pose |
The thresholds are calibrated on observed outcomes, not chosen first: tv remote
1.8° failed on 100 % of calls; fan decoration 26° and projector screen 16°
both lifted, the latter to 0.02 m.
So anchor size is not a general problem — it is loft's problem. A two-stage
"find the anchor first" search would serve 5 % of relations and is not worth
building; resolve_relation already works on the 84 %.
Three caveats: this is measured from the start pose, so the 2–5 % band improves
by driving; it covers 7 of 15 scenes because the rest have no exported ground
truth; and the matcher is crude — 6 phrases matched nothing, and its first
version resolved "the TV remote" to the 1.02 m tv, deleting the very case the
script exists to count. That error mode is one-directional, so the true figure
is likely below 84 %.
Corrected during the session¶
Four claims made here and then measured false. Recording them because the pattern — an explanation that fits, asserted before it was checked — is the recurring failure mode of this work, not any one bug.
| claim | what measurement said |
|---|---|
| "we left 0.42 m on the table at the japanese_room lantern" | the proxy was a z-band over the dataset cloud; the real terrain floor is 1.16 m and we reached 1.06 m |
| "routing exploration through the converter throttles it" | the legal set reached 5.16 m; the model's heading was blocked |
| "the model switched from the coffee table to the dining table" | it said dining table in every step of every run; a subordinate clause had been quoted as the conclusion |
| "loft fails because objects are not in the first frame" | all four failures were small-or-distant targets and ranging errors, none a visibility failure |
State¶
Committed as dad76b3 on feature/perception-replay-harness — 30 files, not
pushed. Everything after that commit (binding, constraints, exploration, v5,
anchor_size.py) is uncommitted.
Open¶
gate: ten questions, parsed but not enforced.- Ground truth for the other eight official scenes; half the question set cannot currently be scored offline.
- Repeat runs per scene. Every comparison above rests on one or two samples
against a spread that is 0.45 m wide on
chinese_room. - Multi-constraint instructions — go near A, then take the path near B to C — are not implemented; the loop drives to one object and stops.
NEAR_M,JUMP_M,KEEPOUT_M,MAX_EXPLORE_Mare calibrated or bounded, not fitted.