Skip to content

Task Reports

Development progress is tracked through numbered tasks. Each report lives in this directory (docs/tasks/) alongside this index.

Task numbers are not unique — several rounds ran in parallel on different branches and reused a number, so 7, 8, 9 and 10 each have more than one report. Every row below links to a distinct document.

Task Title Status
1 Consolidate input/output format Done
2 End-to-end Docker integration Done
3 VLA-3D generated dataset pipeline Done
4 Integrate Qwen3.5 VL for testing Superseded by 17
5 VLM tick logger Done
6 Documentation webpage In progress
7 Data coupling — coverage trajectory generation Done
7 Data sample visualization script Done
7 Evaluation pipeline integration Done
8 Frontier-based room exploration Superseded by 12 / 15
8 Ref dataset ambiguity audit & label-vocab fix Done
9 Expand ref corpus & match official phrasing distribution Done
9 Gemini-backed responder for Task 1 and Task 2 Superseded by 14
10 Offline Gemini evaluation pipeline Done
10 Scene representation Done
11 Perception responder Done
12 Consolidate exploration + scene-building pipeline Done
13 Merge perception enhancements into the pipeline Done
14 Consolidate submission pipeline into the scene_gemini responder Done — live smoke test pending
15 Retire the unreachable Phase-A exploration stack Done
16 Move task reports into docs/tasks/ Done
17 Retire the qwen responder Done
18 Deduplicate the responder factory Done
19 Remove the near-relation subsystem Done
20 Exploration benchmarking across Unity scenes Done
21 Offline perception replay harness & 3D lift repair Done — dedup pending
22 Multi-scene perception benchmark & the median-box fix Done
23 Where recall actually dies, measured stage by stage Done
24 What the boxes are actually missing Done
25 A viewer for the perception map, and what it found Done
26 Can a VLM ground a referring expression, measured Done
27 From a box to a waypoint, and two things measurement changed Done — superseded on blind lifts by 28
28 The approach loop, driven Done — superseded on the platform clamp by 29
29 The converter is not a clamp, and we can predict it Done
30 Four scenes, and the bugs only driving finds Done — loft unsolved, gate not enforced
31 Ground truth for all fifteen scenes, and what it refuted Done — KEEPOUT_M refuted, not yet redesigned
32 Who splits the instruction, and where keep-outs live Done — model parser primary, 30/30 drive order; executor next
33 The ordered plan executor Driven — chinese_room 2/2 against GT
34 Driving a passage, and four ways of faking it Done — 4 passage bugs fixed; studio crosses the gap at x=+3.46
35 Leaving the room the robot started in Driven — home_building_2 q1 2/2; multi-room search still unsolved
36 The leg boundary that erased the robot's momentum Fixed offline — 3/3 replayed reversals gone; not yet driven
37 The keep-out that caused the violation Fixed offline — livingroom_2 q5 replay no longer crosses the gate; not yet driven
38 What an earlier leg already saw Built — sightings carried across legs; not yet driven
39 The curve that was mistaken for a circle Driven — cr_0811_03 leg 2 reaches the painting (2.82 m → 0.83 m, 0.51 m from the reference stop); leg 1 still binds the wrong plant, and the model's own close-range prose names the support correctly with nothing reading it
40 Wiring Gemini, and the declaration that was not true Built — --backend gemini had never completed a call; client lifetime, dead default model, output budget and a false coord_space all fixed; one call per model verified, not driven
41 The comparison that quietly dropped the answer Built — a partial or undecided comparison is demoted, not passed off as a winner or thrown away; replayed 15 GT legs: 1 better (jr_0812_04 3.93 m → 0.64 m), 0 worse; jr_0812_01 still needs the blind-cone margin
42 The pose that was older than the frame Done — audit of PR #28 against the drive loop: extrinsic already ours, Capture had the same unmatched pose/frame pairing (107 ms) and now takes the pose after the image; measured skew was −0.11° ± 0.33, so hardening not repair
43 The box that was the window, and the verdict that was never written Built — o_1_0814_02 drove at the window it was told the cooler was near, because feature_box_2d carried the anchor and the loop preferred it unconditionally; refused now when its lift disagrees with the target box's by more than dominant_cluster's 0.35 m. Replayed 490 steps: 105 boxes dropped, 92 kept, new lift closer to the model's own estimate 29–18. Ctx.record also deferred, so arrived/stopped finally reach steps.jsonl
44 The ceiling that counted the thinking Built — a grounding call spent all 4096 output tokens and wrote no brace: claude-opus-5 thinks by default and thinking is billed against max_tokens. Now XIAO_HEI_CLAUDE_MAX_TOKENS, default 16000 (16000 not higher: non-streaming HTTP timeout). Reasoned from the Opus 5 default and output_tokens == ceiling, not re-run — no API key in that session
45 What the robot is told about where it has been Built, not driven — 123 paired API calls show the visited block does reach the model (removing it beats re-rolling the same prompt, p = 0.016) but its advantage on "prefer somewhere it has not stood" is 0.06 m against a 0.06 m noise floor; --visited {prose,bearing,xy,off} added, prose byte-identical to before
46 The arrival the model said had not happened Built, not driven — o_2_0814_02 reported 3/3 with all three legs stopped at a binding the model put 2.8-3.6x further out, one of them 6.5 m short of a door behind glass; target_state == "far" now vetoes arrival at all four paths. Refuses 8 of 39 recorded arrivals: 6 check out against the scan, 2 may be false. A rule on distance_m metres was measured and rejected — it vetoes two correct legs to catch one bad one
52 Porting the drive loop into the submission's ai_module Built, not driven — the loop ran the wrong way round (laptop → ssh → docker exec); now a ROS node inside iros2026_ai_module. Robot's four methods reimplemented in process as RobotNode, so the eleven grounding/nav files are byte-identical and sync_ai_module.sh --check proves it. /camera/image raw replaces the compressed topic (allowed-list), /way_point_reached dropped, budget measured from process start, sensor QoS depth 5 → 1 (a persistent node would have snapshotted a stale frame). Package kept as dummy_vlm/dummyVLM so the graders' startup script still starts us. Two bugs found by moving it: decompose wrote its cache unguarded to a relative path on the answer path, and nothing caught an exception at the top of a run
53 Scoring instruction-following on the challenge's own rubric Done — proxy validated against the organisers' own reference answers (5.97/6, 29 of 30 at full marks); negative controls added, which found and fixed the order anchor; the 25 recorded runs score 4.56/6 but cover only 9 of the 30 official questions, so that number is a development log and not a system score
54 Falsifying a binding with the motion already paid for Measured offline — of four candidate cross-step tests, three fail and T1 (predicted range vs odometry) catches 44% of wrong bindings at 5% false alarm, rising to 69% at 22.8x separation once it is allowed to abstain where its geometry gives it no leverage. Not yet wired into the loop
55 A verifier must not be told what it is verifying Measured — a leg running out of time was crashing whole questions (Ctx.left() went negative into subprocess.run(timeout=)); plan.json now survives any crash or Ctrl-C. Then the redirect: T1's branch is 2.2% of steps while no measurement at all is 21.3%, so the semantic verifier is gated on the second. Its first form refuted right bindings more often than wrong ones and said holds on 83% of bindings beyond 5 m against 35% within 5 m — it agreed more the less it could see, because the prompt named the hypothesis. Blinding the looking call takes false confirmation of a wrong binding from 5/10 to 0/10 (p≈0.03). Also corrected two drafted claims against the data: run-to-run spread is a minority of unstable questions (3 of 4 repeats identical), and a leg's own verdict is not the rubric's — jr5_p1 scores 6.00/6 while reporting itself 2.40 m short. Live arm not yet driven
56 The scorer was reading the noun and ignoring the clause Found by drawing the runs, not by testing them: head_label reduces "the potted plant on the table" to potted plant, so any instance of the noun satisfied a destination and the relative clause — the whole job — was never tested. 72% of credited destinations sit in scenes with rival instances and 15% were credited on a different instance than the organisers' own reference path went to, worst case 8.2 m away. reference_pins pins each ambiguous GOTO to the instance their reference trajectory approaches most closely; --strict-instance scores against it. Corpus mean 4.29 → 3.77/6, full marks 25→19. The conclusion it changed: under the loose reading misses pile against the tolerance (44% near, tail ends at 4.2 m) and read as a platform ceiling; strictly, half are in a tail out to 8.2 m and are grounding failures — ours. The loose reading relocated the blame, not just the score
57 Counting is a different question than pointing The numerical responder, built and gated. Read as a set, 11 of the 15 released questions are anchor-local — one piece of furniture, count what is on it — so the task reuses run_goto almost whole. New: a blind counting call (never told the running total, per TASK 55), a lidar merge so the answer is the number of clusters and never the sum of the views, and an answer key computed from VLA-3D (which does carry colour: home_building_2 is 2 maroon pillows). Gate, replayed over 32 recorded views at zero simulator cost: 38% exact against a 27% constant-guess baseline — 71% on the five scenes whose answers are 1-3 and 11% on the six whose answers run to 8, with 0/13 on full views of the hard ones. The first report said 71% and was wrong twice over: it tested only small counts, and visible_from read across as centred when it is u/FACE_SIZE, discarding everything right of each face's centre — the merge was built for exactly that and could not act, because the scanner placed 0 of those boxes (28/31 overall; every failure is a cup). Also found, unrelated: verify_binding.py was missing from the ai_module sync list while approach_loop imported it at module scope — the submission would have died at launch, taking all 36 instruction-following points
21 Reject mask spill in the 3D lift Done
22 Same-label duplicate suppression in ObjectMap Done
23 Capture perception inputs from real navigation Scripts done — sweep not yet run
24 Audit of un-synchronised gating logic C resolved; A superseded (see note); B, E, F open
25 Merge detections split by the panorama face seams Done
26 Scan-accumulator keyframe sweep & side-by-side dump comparison Dumps, viewer & GIF tool done; scored sweep not run
27 Verifying the 2D-detection → 3D-lift angular convention Done — extrinsic bug fixed + image/pose timestamp-matched in LatestCache
28 Swap YOLO-World for OWLv2 & drop the tick to 1 Hz OWLv2 measured & reverted (over-produced + too slow); kept 1 Hz + SAM-Large
29 Watching the carpet fragment, frame by frame Done — cross-frame fusion animator; seam-merge overlay left as follow-up
30 Raising the detection floor to 0.6 and gating the seam merge Done — 0.6 default stack-wide, re-dumped; carpet 13→9; e2e re-score pending
31 GT footprints ignored heading, and a flat-object box finding Done — GT heading fixed in dump_debug; flat-object box fix proposed, not implemented
32 A flat-aware, gap-based merge gate for ObjectMap Done — carpet 9→2 on arabic_room; multi-scene sweep pending
36 Viewpoint redundancy, the dwell bias, and a flat-object extent estimator Done — flat max-extent + union-midpoint centre (carpet coverage 20→56%, 64→87%); novelty gate default ON 0.3 m; multi-scene sweep pending
37 Detection confidence audit, score/SAM thresholds, and the live-threshold tool Done — score 0.6→0.4 + SAM 0.8 gate + keyframes→2 (perception mAP@1.0 +43%); controlled e2e +29% on captures_nav; live-threshold viz tool; live-explore re-validation pending
38 VLM-based navigation waypoint proposer (Opus 5) Scaffold done — drop-in nav_vlm strategy, event-triggered proposer, reachability snapping + failure feedback; unit-tested against a fake engine; live validation pending Anthropic key
39 Question-directed VLM navigation to the referenced object Scaffold done + both Claude paths live-validated on Opus 5 — nav_task1 explorer drives to the object named in an object_reference question (Claude on reach/skip), then scene_claude responder dumps the scene graph to Claude to answer; 29 mode tests pass; full sim run pending
47 Periodic exploration snapshots instead of end-of-sweep only Done
48 Artefact paths keyed on scene and strategy Done
49 An eight-minute wall-clock cutoff for exploration Done
50 Why both explorers stall, and which one to keep Frontier picked; 6 fixes landed + 13 tests — validated by the exp1 sweep: frontier 2516 m2 over 13 scenes, +59% median coverage
51 Three exploration changes that were tried and reverted Reverted — none beat exp1. Kept as the record of what not to retry
52 A 100-label detector prior covering all 15 scenes Done — --max-labels cap (ranked by cross-scene generality); 100 labels, 81% GT coverage, every scene 68–96%

The task definitions these reports answer to stay at the repo root in TASK.md.

Backlog

Diagnosed but unfixed gaps, with the evidence attached, live in backlog.md.

Design documents

Detailed technical documents are in the docs/ directory:

Historical design records

Kept for their recorded rationale; they describe code that no longer exists.