TASK 24 — Audit of un-synchronised gating logic¶
Why¶
The stack answers the question "is this new?" in eight different places, with eight independently-chosen thresholds, and nothing reconciles them. Six of those are pose-based, so "the robot is somewhere new" is currently defined as 0.25 m, 0.3 m, 1.5 m, 2.0 m, 3.0 m, and a 0.92-valued nav signal — depending on which layer is asking.
This is an audit, not a change. Nothing here is fixed yet; each item records what the logic is, what breaks because of it, and what the evidence says. The point is that picking any of these up later does not mean re-deriving the diagnosis.
Explicitly out of scope: the tick itself (0.5 s, XIAO_HEI_VLM_TICK_HZ) and
sensor pairing (LatestCache.snapshot). Both live inside tick() and are
deliberately left unchanged. See "Deferred" at the end for why, and for the one
measurement worth taking there.
The inventory¶
| # | layer | trigger | value | source |
|---|---|---|---|---|
| 3 | scan keyframe | ~~motion since last keyframe~~ → every tick | window 10 ticks, voxel 0.05 | perception/scan_accumulator.py |
| 4 | viewpoint node | distance from every existing node | ≥2.0 m | scene/representation.py:344 |
| 5 | object merge (add_object) |
same-label proximity | 1.5 m | scene/representation.py:116 |
| 6 | object merge (ObjectMap) |
IoU or centre dist | IoU 0.3 / 0.4 m; NMS: IoU 0.5, dist 0.4 + gap 0.05 | perception/object_map.py:37-59 |
| 7 | lift acceptance | inlier count | min_inliers=10, cluster voxel 0.10 |
perception/lifter.py:54,78 |
| 8 | waypoint reached (explorer) | distance | 0.3 m | exploration/_frontier.py:66 |
| 9 | waypoint reached (nav signal) | a different signal entirely | 0.92 | app/main.py:483 |
| 10 | explorer give-up | time / count | stuck 12 s, max skips 20, max wp dist 1.5 m | app/main.py:273-274 |
A — #3 vs #4: two motion gates that cannot be related to each other¶
Superseded, not resolved. #3's motion gate has since been deleted — the accumulator now commits a keyframe on every tick, so its window is keyed on ticks rather than travel. That removes the mismatch described below by removing one of the two gates, but it does not unify the layers: #4 still keys on distance. It also introduced two costs measured after the change:
- Coverage is evicted by time.
max_keyframes=10at 2 Hz means the buffer turns over after 5 s. Simulated over 5 distinct regions: driving past all five leaves 5 regions in the merged cloud; 10 stationary ticks leave 1. On the TASK 23 captures — 66% ofarabic_roomframes below 0.05 m/s,chinese_roomnever moving — densification now switches itself off exactly when the robot is stalled.- Every tick pays the merge.
update()used to return the cache untouched off-keyframe. It now vstacks and voxel-downsamples on every call: 218 ms per tick on a 20k-point sweep, against a 500 ms budget, on top of detect + lift + the sidecar round-trip.Keying eviction on travel rather than ticks would keep the simplification and recover both properties. The measurements below predate the change and describe the gate as it was.
ScanAccumulator._should_keyframe stores a sweep when the robot has moved
≥0.25 m or turned ≥15° since the last stored keyframe — a chained delta.
SceneRepresentation._maybe_add_viewpoint adds a node when the pose is ≥2.0 m
from every existing node — a global minimum-distance test. Different
reference frames, different predicates, and rotation counts for one and not the
other.
They are therefore not convertible. Measured over the TASK 23 navigation captures (15 scenes, 5487 ticks):
| scene | ticks | scan keyframes (#3) | viewpoint nodes (#4, 2 m) | #4's rule at 0.25 m |
|---|---|---|---|---|
| arabic_room | 316 | 147 | 3 | 37 |
| chinese_room | 228 | 1 | 1 | 1 |
| home_building_1 | 767 | 313 | 31 | 241 |
| home_building_2 | 1008 | 696 | 11 | 117 |
| hotel_room_1 | 377 | 211 | 8 | 82 |
| hotel_room_2 | 207 | 59 | 2 | 20 |
| japanese_room | 166 | 64 | 5 | 44 |
| livingroom_1 | 297 | 119 | 5 | 50 |
| livingroom_2 | 378 | 187 | 6 | 57 |
| livingroom_3 | 364 | 128 | 7 | 52 |
| livingroom_4 | 227 | 104 | 6 | 43 |
| loft | 280 | 31 | 2 | 14 |
| office_1 | 328 | 181 | 10 | 88 |
| office_2 | 247 | 91 | 5 | 37 |
| studio | 297 | 118 | 5 | 51 |
| total | 5487 | 2450 | 107 | 934 |
The last column is the key evidence: run #4's predicate at #3's threshold and you get 934, not 2450 — a 2.6× gap from the predicate alone, before any threshold disagreement. So there is no fixed ratio between "sweeps in the accumulator" and "viewpoints in the graph", and neither can be inferred from the other.
Proposed direction: collapse both into one keypose. A single tracker owns the
0.25 m / 15° chained test and emits a keypose_id per tick; ScanAccumulator
consumes it instead of re-deriving motion, and _maybe_add_viewpoint fires
exactly when a new keypose is emitted. Three things fall out:
- the accumulator window becomes literally "the last 10 viewpoints";
observing_viewpoint_idsbecomes real provenance at ~0.25 m granularity;- one knob (
XIAO_HEI_SCAN_MIN_MOVE_M) controls both, so they cannot drift.
Cost: 107 → ~2450 viewpoint nodes across 15 scenes (~163/scene, tens of KB of
JSON). Blocker to handle first: RoomNode.best_image updates on every new
viewpoint node (representation.py:354-358). At keypose density that becomes
per-keyframe image churn, so it needs its own explicit "best coverage" rule
before the layers are merged.
B — #4: the viewpoint edge points at the newest node, not the nearest¶
_current_viewpoint_id() (representation.py:324-342) returns
self._viewpoints[-1].tick_id. Both object paths use it — add_object at
line 186 and sync_from_object_map at line 236.
Because #4's gate only requires 2 m from some node, the robot can sit 0.5 m from vp0 while every detection is attributed to vp7 several metres away. Concretely: nodes laid at (0,0), (2,0), (4,0), then the robot returns to (0.5,0). No new node — it is 0.5 m from vp0 — but the edge still records vp2, 3.5 m off. On the TASK 23 trajectories, which are short and heavily back-and-forth, this is the common case rather than the corner case.
observing_viewpoint_ids is therefore not currently trustworthy as geometry.
Secondary defect: dedup only compares against the last entry
(if not vp_ids or vp_ids[-1] != vp_id, lines 205 and 237). An object seen from
vp0 → vp1 → vp0 ends up with [vp0, vp1, vp0]. So len(observing_viewpoint_ids)
counts visits, not distinct viewpoints; anything reading it as a view count
needs a set().
Proposed direction: return the nearest viewpoint, not the last. ~5 lines, no dependency on A, and it makes the existing edges correct. Do this one first.
C — #5 vs #6: not two merge rules, one flag with inconsistent defaults¶
Resolved.
add_objectandmerge_radiusare deleted;ObjectMapfusion is unconditional andXIAO_HEI_OBJECT_MAPis gone from the code, the four config files and the docs. The offline constructors ingemini/batch.py(which had no point clouds to fuse) now build a merge-disabledObjectMapfrom their boxes and sync it in, so there is exactly one path into the object layer. The diagnosis below is kept for its rationale.
These are mutually exclusive implementations of the same step, not layers.
_inject_visible branches per detection (responder.py:311-317): when
self._object_map is not None the lifted cloud goes to object_map.add(...)
and continue skips scene.add_object entirely; line 331 then replaces the
whole object layer via sync_from_object_map. merge_radius=1.5 and the
IoU-0.3/0.4 m rule never both apply.
add_object is the original — same label + close centre, higher confidence
wins, keeping one frame's median and no 3D box. ObjectMap is the
replacement: it unions the point clouds so the AABB converges, plus NMS and
wall-sheet rejection. TASK 13 added it as opt-in so the pipeline would "behave as
before" by default, and kept add_object as the arm of an on-sim A/B that
TASK 13's own notes record as never run.
The duplication is not the problem. The problem is that the flag defaults differently per entry point:
| entry point | XIAO_HEI_OBJECT_MAP |
|---|---|
docker/compose.yml:72 |
${...:-} → off |
docker/compose_scene_gemini.yml:125 |
${...:-} → off |
docker/compose.eval.yml:59 |
${...:-1} → on |
scripts/run_scene_vla3d_eval.sh:175 |
${...:-1} → on |
perception_benchmark/replay_score.py:88 |
hardcoded ObjectMap(...) — no opt-out |
So every perception number on record — TASK 20–22, the TASK 23 navigation
scores, the scan-accumulator A/B — measures the ObjectMap path, while the
default compose.yml runs the add_object path. replay_score.py cannot
measure the path the default compose runs even if asked.
The consequence is concrete: on the add_object path the responder passes
bbox_min=None, bbox_max=None (responder.py:325-326). Under compose.yml's
default the submission's scene graph ships objects with no 3D boxes at all.
Proposed direction: flip the default on and delete add_object +
merge_radius. ObjectMap is strictly more capable and is what every
measurement was taken on; a fallback no benchmark can score is exactly how this
divergence arose. Minimum acceptable fix: make the four defaults agree.
Open question, must be answered first: which compose file the challenge submission actually launches. Both candidates currently default to off.
D — #7: thresholds in object space, not pose space¶
min_inliers=10, cluster_voxel=0.10, and #6's IoU/distance values are tuned
against measured data (TASK 21, TASK 22) and are not in the same units as the
pose deltas above.
No action. The un-synchronisation here is intentional and should stay that way. Recorded so it is not swept into a later "unify the thresholds" change.
One caveat carried from TASK 20: the min_inliers sweep was never re-run with
the scan accumulator enabled, so 10 is tuned against single-sweep density while
the live stack lifts against an accumulated cloud. That is a measurement gap,
not a synchronisation one — it stays in backlog.md.
E — #8 vs #9: two unreconciled definitions of "reached"¶
FrontierExplorer has waypoint_reach_dist=0.3 (_frontier.py:66), a distance
in metres. app/main.py:483 has _WP_REACHED_THRESHOLD = 0.92 applied to a
Float32 on /way_point_reached — a different signal in different units, with
the comment "nav stack settles between 0.25-0.90 m depending on obstacles".
Nothing reconciles them, and both can gate progress.
F — #10: the give-up constants are hardcoded and are firing everywhere¶
stuck_timeout_s=12.0 and max_consecutive_skips=20 are literals at
app/main.py:273-274 (overriding the class defaults of 30.0 and 3). Neither is
env-settable, so TIMEOUT in the capture sweep cannot influence them.
Every TASK 23 run ended early:
| scene | visited | skipped | reason |
|---|---|---|---|
| home_building_1 | 26 | 37 | max_consecutive_skips |
| livingroom_1 | 9 | 20 | max_consecutive_skips |
| livingroom_3 | 8 | 25 | max_consecutive_skips |
| office_2 | 6 | 15 | no_frontiers |
| hotel_room_1 | 5 | 26 | no_frontiers |
| japanese_room | 5 | 11 | no_frontiers |
| studio | 5 | 20 | no_frontiers |
| loft | 4 | 22 | max_consecutive_skips |
| arabic_room | 3 | 24 | max_consecutive_skips |
| livingroom_2 | 3 | 27 | max_consecutive_skips |
| office_1 | 3 | 21 | no_frontiers |
| livingroom_4 | 2 | 18 | no_frontiers |
| chinese_room | 1 | 20 | max_consecutive_skips |
| hotel_room_2 | 0 | 18 | no_frontiers |
| home_building_2 | — | — | no DONE line — truncated at the 510 s cap |
skipped exceeds visited in every scene. chinese_room recorded a path
length of 0.0 m over 228 ticks; hotel_room_2 visited zero waypoints. This
is the single largest confound on every navigation-capture number, and #8/#9 are
a plausible contributor: if the two "reached" tests disagree, the explorer
re-targets a waypoint the nav stack considers reached, burning a skip.
Proposed direction: pull the WP_SKIP lines from
exploration_logs/<scene>/exploration.log and check whether the frontier
targets land in untraversable space or merely time out short. Treat as an
exploration bug, separate from the perception work.
Suggested order¶
- ~~C — resolve the
XIAO_HEI_OBJECT_MAPdefault.~~ Done — fusion is now unconditional, so the submission always ships 3D boxes. - B — nearest-not-newest viewpoint edge. Small, self-contained, fixes a real defect in data already collected.
- F/E — diagnose the explorer stall. Blocks meaningful scoring of the navigation captures.
- A — the keypose merge. Largest change, and worth doing only once the
best_imagequestion above is settled.
Deferred: the two tick-internal items¶
#1, the tick (0.5 s). Unchanged by decision. Object fusion runs
unconditionally every tick — _inject_visible skips only when the snapshot
lacks an image, pose or scan. That cadence is the reference the TASK 23
recorder reproduces, and changing it would invalidate the captures.
#2, sensor pairing. cache.snapshot() is the first statement of tick()
(app/main.py:505), so the perception stack only ever sees one frozen
VLMInput — image, pose and scan are guaranteed same-generation for the whole
tick, and the sidecar round-trip cannot skew them apart mid-tick. But
snapshot() (sync/latest_cache.py:79-92) merely copies whatever is in each
slot under a lock. There is no timestamp comparison, no tolerance window, no
staleness rejection — it is an atomic read, not a synchroniser. With
/camera/image at ~10 Hz and the scans at ~5 Hz, channel skew of ~200 ms
(≈20 cm of ego-motion at 1 m/s) is unmodelled.
Two facts for whenever this is picked up:
rclpy.spin(node)(main.py:679) uses the default single-threaded executor, so no callback runs whiletick()executes. Channel ages are fixed at snapshot time. The flip side: a tick longer than ~500 ms stops all draining, and the subscriber QoS isdepth=5, RELIABLE(adapters/ros/subscribers.py:69-73), so that queue will overflow.- Every message already carries its stamp —
_header_from_ros(msg.header)is preserved onImageFrame,LidarScan,TerrainMapandOdomPose. The data needed to measure skew exists and is unused.
The non-invasive next step, if wanted: have snapshot() also report per-channel
age (tick_time - header.stamp) and log it. Measurement first, no behaviour
change, then decide from real numbers whether a tolerance gate is warranted.
Related¶
- TASK 13 — introduced
ObjectMapas opt-in; its unrun A/B is the origin of C. - TASK 20–22 — every number they report was measured on the ObjectMap path.
- TASK 23 — source of the navigation captures all measurements above draw on.
docs/tasks/backlog.md— B2, B3 and themin_inlierssweep remain open and are unaffected by this audit.