TASK 41 — The comparison that quietly dropped the answer¶
runs/jr_0812_01, "Go to the lantern closest to the fan decoration, then take
the path near the wardrobe doors to the flowers on the display ledge".
| leg | reported | truth |
|---|---|---|
| 1, the lantern | arrived |
4.25 m from the right lantern |
| 2, the wardrobe doors | passed |
correct |
| 3, the flowers | arrived, circled back (2.06 m) |
1.61 m from flowers id6; the binding was 0.66 m from it |
japanese_room has three lanterns. lantern id61 at (+1.40, -0.52) is
0.63 m from fan decoration id62; the other two are 6.5 m away in the
tatami hall. The answer is id61.
Not the grounding, and not the anchor¶
The anchor was right on every call:
| true fan bearing | model's anchor | error | |
|---|---|---|---|
| step 1 | az −45.6°, el +18.2° | az −43.3°, el +15.8° | 2.3° / 2.4° |
| step 2 | az +177.2°, el +7.0° | az +176.8°, el +7.5° | 0.4° / 0.5° |
And at step 2 the model listed the correct lantern:
model "black floor lantern beside the low stand under the fan" az -174.6 el -8.3
truth lantern id61 az -174.8 el -7.8
0.2° and 0.5°. The identification was not the problem.
resolve_relation dropped it and declared a winner anyway¶
boxes() lifted each candidate to a map position and kept the ones that
lifted:
xy = _lift_xy(px, i, scan_cam, pose)
if xy is not None:
out.append(...) # and if it is None, nothing is recorded
The floor lantern sits directly behind the robot at azimuth −174.6°, where the
scanner's elevation floor is +9.4°. At −8.3° it is 17.7° under it, so the lift
refused it. len(cb) >= 2 still held over the other three, and the comparison
returned "large slat lantern hanging from the coffered ceiling of the tatami
hall = 3.74 m" — 3.98 m from the answer — with a rationale that reads as
authoritative in the log.
Fewer than two candidates was the only incompleteness the function checked for. Fewer than all is the other one, and it is the dangerous one, because it still produces a number.
Replaying every relational reply in leg 1 through the new code:
| step | candidates reported | compared | dropped |
|---|---|---|---|
| 1 | 4 | 4 | — (the answer was never listed) |
| 2 | 4 | 3 | the answer |
| 3 | 5 | 3 | two ledge lanterns |
| 4 | 5 | 3 | two ledge lanterns |
Across all recorded runs, 243 comparative replies:
| resolved and complete | 120 (49%) |
| resolved but partial | 20 (8%) — previously indistinguishable from complete |
| unresolvable, falls back to the model's pick | 103 (42%) |
Median one candidate dropped, at most two.
What changed¶
resolve_relation now returns a Resolved dataclass carrying complete and
missed, and still iterates as (box, image_index, why) so both existing call
sites destructure unchanged.
- The reason string names what was left out (
"; NOT compared: ..."), so a log reader is not misled the way this one was. - The winner is demoted, not discarded.
verifiedis untouched — an incomplete comparison is still the best measured evidence there is, and refusing to bind on it would keep whatever wrong binding came before, which is what this run already did. What is withdrawn ismeasured, the licence that lets a reading overrule an earlier binding at any distance. A comparison missing a candidate has not earned that, because the candidate it could not lift is exactly the one that might have won. - The dropped candidate's bearing is kept and fed back. A candidate the
scanner cannot reach is usually behind the robot, and that is a direction to
turn rather than a thing to forget. It goes into
ctx.visitedas a lead: "a possible 'black floor lantern...' was seen at bearing −175° but could not be measured from here".
Undecided is not unmeasurable¶
runs/jr_0812_04, the same lantern phrase, found the second half of the same
bug. resolve_relation returned None whenever fewer than two candidates
lifted, and the caller reads None as "unverified", which makes bind_target
keep whatever binding it already has.
At steps 3 and 4 the anchor lifted to 0.11 m and 0.02 m of the true fan
decoration, and exactly one candidate lifted — the right lantern, 0.57 m and
0.64 m from it. Both calls produced a committed waypoint 0.05 m and 0.04 m
from the truth. Both were discarded for a binding carried from step 2 that
sat 3.93 m away, and the leg then reported arrived, within standoff because
the robot happened to stop 0.25 m from that wrong binding. The two candidates
that failed to lift were the tokonoma ledge lanterns, behind the robot in the
blind cone — the wrong ones.
One survivor out of several is an undecided comparison, not an unmeasurable
one, and it now returns a Resolved marked incomplete: the same demotion as a
partial comparison, because it is one. The anchor-never-lifted case is
untouched and still returns None — that is the loft counterexample the
branch was written for, where the phrase carries no evidence at all.
Except on farthest_from, where the survivor argues against itself¶
The first driven run after this change, runs/ar_0812_03, was made worse by it,
and the reason generalises. A lift fails because the object is far, or
occluded, or under the scanner's floor — so "the only candidate that lifted" is
systematically the near one, and farthest_from is precisely the question whose
answer is the far one.
Leg 1 asked for "the potted plant furthest from the hookah". One of two
candidates lifted; it was the nearest plant of five, 8.59 m from the answer.
Binding it locked the leg out of the two steps that followed, whose committed
waypoints sat 0.37 m and 0.55 m from the truth and were refused by JUMP_M
because a binding was already held. Before the change no binding was made and
those two steps were free to drive at the right plant.
jr_0812_04, the case the rescue was written for, is closest_to — where the
same bias points at the answer. So the rescue applies to closest_to and
between and returns None on farthest_from, which is what the code did
before.
Replayed through the real bind_target chain, the fix reaches the right
lantern by the route already in the code:
step 3 jump 4.49 m > JUMP_M -> refused, recorded in `pending`
step 4 two readings within 1.0 m of each other and far from the binding
-> corroborated -> re-bound (-1.83,+1.73) 3.93 m out
-> (+1.93,-0.87) 0.64 m out
and the robot is then 4.1 m from its binding, so it keeps driving instead of declaring arrival.
Regression over every leg with ground truth — 15 legs across chinese_room,
japanese_room and livingroom_2, replaying the whole binding chain both ways:
| better | 2 (jr_0812_04 leg 1, 3.93 m → 0.64 m; jr_0812_05 leg 1, 6.29 m → 0.04 m) |
| worse | 0 |
| unchanged | 14 |
ar_0812_03 leg 1 never reaches a final binding under either version, so it
falls outside that table; checked separately, it is back to the pre-change
behaviour of not binding at all.
641 tests pass, 7 new — five of them replaying recorded replies against recorded scans, so the fixtures cannot drift from what happened.
This does not fix jr_0812_01¶
Stated plainly because the run is the reason the task exists. Step 3 found the right lantern and the geometry picked it — "small black slat lantern/basket on the floor beside the low stand under the fan panel = 0.77 m", against a ground truth of 0.63 m — and it was thrown away by a different test:
winner bearing az +113.8° el -4.38°
scanner floor there -4.16°
-> 0.23° below -> blind -> lift not trusted -> binding carried
The lift that was refused would have landed at (+1.68, −0.91), 0.48 m from
the truth, against the 3.98 m the leg kept.
COVERAGE_FLOOR_DEG is not wrong — rebuilt in the sensor frame from 498 scans
across 74 runs it matches the table to within 1° at every bin. It is a 0.5th
percentile, so 0.5% of returns lie below it by construction, and the
scene-to-scene spread is about 2°. Treating it as a knife-edge is the bug. Of
69 blind rejections in the recorded runs the median sits 9.8° below the floor
and the deepest 37.3°, but 11 (16%) are within 3°, and 10 of those 11 had a
lift available:
| below floor | run | |
|---|---|---|
| 0.23° | jr_0812_01 s3 |
this one |
| 0.23° | cr5_conv s3 |
|
| 0.38° | cr_0811_03 s6 |
|
| 0.44° | hm2_v6_2 s3 |
|
| 0.77° | lr_2_0811_07 s8 |
|
| 2.5–2.8° | lr1_0811_0{1,2,3} s1 |
the first step of three runs |
A margin is the obvious fix and is not made here. TASK 27's counterexample is real — a target 17° into the blind cone lifted 4.70 m against a true 2.9 m because the cone widened upward onto the wall above it — so the margin should be chosen after computing the lift error for all eleven, not from this one case.