TASK 23 — Capture perception inputs from real navigation¶
Why¶
Every perception number to date (TASK 20–22, the clustering benchmark, the 2D
detector audit) was measured on perception_benchmark/captures/, which
capture_viewpoints.py produced by driving the robot to the geometric k-cover
viewpoints from viewgen.py. That dataset answers "how good could perception
be given near-ideal coverage?" — 88–100% of scoreable objects are guaranteed
observable from at least one captured pose, by construction.
It does not answer what the submission stack actually sees. Live, the robot follows the frontier explorer, which optimises for map coverage, not object coverage: it stops where frontiers are, at whatever range and angle that implies, and it never revisits a pose for a second look. The scene graph the challenge is scored on is built from that stream.
So the two datasets can disagree in both directions — a viewpoint-set object may never be seen live, and the live run produces many more, closer, partially redundant views. Task 23 builds the second dataset so the gap can be measured instead of assumed.
What was built¶
perception_benchmark/record_navigation.py¶
A passive rclpy recorder. It publishes nothing — the ai_module container
drives with the real FrontierExplorer, and this node only watches
/state_estimation and the sensor topics.
It reproduces the live ingest cadence, which is the point. In app/main.py
the node ticks at XIAO_HEI_VLM_TICK_HZ (default 2.0 → every 0.5 s), and every
exploration tick calls responder.ingest() → _inject_visible(), which runs
detect → lift → add_object unconditionally. There is no distance, rotation
or motion-blur gate on that path: a tick is skipped only when the snapshot lacks
an image / pose / scan, and a blurred frame simply returns no detections and
leaves the scene unchanged. So the recorder samples at a fixed --rate-hz
(default 2.0) and applies no quality filter.
The first draft gated keyframes on 0.6 m / 30° of motion with a motion-blur speed gate. That was wrong — it is a capture policy, not the system, and it would have handed the offline replay ~10× fewer, sharper, better-spaced frames than the live stack ever sees, quietly flattering every downstream number. Those gates were then kept as opt-in knobs, and finally removed outright: an option that can only make the dataset diverge from the system is not worth the config surface, and its presence invites exactly the silent divergence the task exists to eliminate.
A warm-up delay went the same way, and was the subtlest of the three. The live
node starts its timer immediately and skips ticks whose snapshot is incomplete
(_inject_visible returns early on a missing image / pose / scan). Spinning for
3 s before the first sample would have suppressed ticks the live stack would
have attempted. The skipped.incomplete counter reproduces that behaviour
faithfully instead — a small non-zero value at startup is correct, not a fault.
--rate-hz now defaults to XIAO_HEI_VLM_TICK_HZ itself, so the recorder tracks
the live tick rate rather than hard-coding a copy of it.
The only frames dropped are ones where the camera has not published since the
last sample. That needs a stalled camera — /camera/image runs ~10 Hz
(docs/architecture.md:68) against a 2 Hz tick — so in a healthy run this never
fires. It is skipped rather than written as a byte-identical copy, and
counted: a non-zero count means the live stack would have re-ingested and
re-fused that image, so it is a real divergence worth seeing. All three skip
counters (incomplete, stale_image, downsampled) land in manifest.json.
An earlier draft paired this with an --allow-duplicate-frames escape hatch to
restore live fidelity. It was removed: at 10 Hz vs 2 Hz the case it guards
against cannot arise outside a fault, so it was a knob bought for nothing. The
counter stays, since that is where the diagnostic value actually was.
The full flag list, and why each survives¶
--scene and --out name the output. --rate-hz is a system parameter
(XIAO_HEI_VLM_TICK_HZ). --stop-file is mechanism, not policy: the live node
ends when the explorer completes, and a passive recorder has to be told. That
leaves exactly one flag with no counterpart in the live stack — --max-seconds,
a disk backstop, since at ~500 MB/min an unattended recorder is a real hazard.
When it fires the manifest records stop_reason: max_seconds, so a truncated
dataset can never be mistaken for a complete one.
There is no motion gate, no blur gate, no warm-up and no frame cap.
The output layout is byte-identical to capture_viewpoints.py's
(vp_XXX/{image,registered_scan,sensor_scan,terrain_map,terrain_map_ext}.npy
+ pose.json + meta.json, + manifest.json), so every offline consumer reads it
unchanged. The wire decoders (image_to_array, cloud_to_array) are imported
from capture_viewpoints.py rather than copied — the intensity-column handling
is subtle enough that a second copy would drift.
manifest.json is rewritten after every frame via temp-file + rename, so an
abrupt kill costs at most the frame in flight.
Disk cost is the real constraint. A frame is ~4.2 MB, 3.7 MB of which is the
raw 640×1920 uint8 image. At 2 Hz that is ~500 MB per minute of exploration per
scene, so a 10-minute scene costs ~5 GB and a 15-scene sweep runs into tens of
GB. RATE_HZ is the knob; check free space first.
perception_benchmark/run_nav_capture.sh¶
Scene sweep. Per scene: extract, bring up a fresh sim + ai_module
(XIAO_HEI_RESPONDER=dummy, so capture needs no GPU and no perception sidecar),
launch the sim, start the recorder inside the sim container, then poll
exploration_logs/<scene>/exploration.log for the explorer's DONE line —
which is the only authoritative end-of-trajectory signal, since the recorder is
passive. On DONE it touches the recorder's stop file, waits for the flush, and
copies the captures to the host.
Resumable (skips scenes with an existing manifest unless FORCE=1), and the
containers are removed on normal exit and on Ctrl-C.
PERCEPTION_CAP_DIR¶
CAP_DIR in the eight offline scripts (replay_score, dump_debug,
dump_cloud, dump_lift_input, detect_2d_eval, dump_detections,
box_quality, debug_viewpoint) now reads PERCEPTION_CAP_DIR, defaulting to
the existing perception_benchmark/captures. This mirrors the
PERCEPTION_DEBUG_DIR knob viz_app.py already had, and means the two datasets
coexist instead of one overwriting the other — the TASK 20–22 numbers stay
reproducible.
Run¶
perception_benchmark/run_nav_capture.sh arabic_room # one scene first
perception_benchmark/run_nav_capture.sh # all 15
export PERCEPTION_CAP_DIR=perception_benchmark/captures_nav
export PERCEPTION_DEBUG_DIR=perception_benchmark/debug_nav
uv run --extra perception python perception_benchmark/replay_score.py --all
uv run --extra perception python perception_benchmark/dump_debug.py --all --out "$PERCEPTION_DEBUG_DIR"
uv run --extra perception python perception_benchmark/dump_cloud.py --all --out "$PERCEPTION_DEBUG_DIR"
uv run --with streamlit --with plotly streamlit run perception_benchmark/viz_app.py
Knobs: RATE_HZ (default 2.0 — match the live tick), MAX_WAYPOINTS (explorer
budget), TIMEOUT (per scene, default 1800 s), FORCE, KEEP_UP.
Verification¶
The sampler is exercised offline with ROS stubbed out (scratchpad
test_gate.py): the 0.5 s period fires at 0.0/0.5 s and not at 0.2/0.4 s; a
stationary robot still yields a frame every tick (the property the first
draft got wrong); an incomplete snapshot is skipped and counted, matching
_inject_visible's own guard; and a stale camera frame is skipped and counted.
Two assertions guard against the mistake this task kept making: a barely-moved
and a fast-moved robot are both still sampled (no motion may ever gate a frame),
and the CLI source is checked to contain none of --min-move, --min-rot,
--warmup-s, --max-frames, --speed-max or --allow-duplicate-frames while
--rate-hz still defaults from XIAO_HEI_VLM_TICK_HZ. All pass.
Not yet run against the sim — that is the next step, and the notes below are what to check when it is.
What to look for once it runs¶
- Frame count vs exploration duration. At 2 Hz the count should be
~2× the seconds between the explorer's
STARTandDONElines. A large shortfall means skipped ticks — read theskippedblock andstop_reasonin the manifest before concluding anything about perception. - That the explorer actually explored. The one real log on disk
(
exploration_logs/studio) endedvisited=0 skipped=20 reason=max_consecutive_skipsafter 261 s. A run like that produces a perfectly valid recording of the robot going nowhere. Checkvisitedbefore scoring. - Recall vs the viewpoint arm. A drop is the real cost of coverage-driven navigation and is the number that matters for the challenge.
- Precision and counting MAE. More, closer, more redundant views push fragmentation up. This is the same axis the scan-accumulator A/B moved (P@1 0.327 → 0.408 with the accumulator off), so the two interact and should be read together.
skipped.stale_image. Should be 0. Non-zero means the camera stalled below the 2 Hz tick, so the live stack was re-ingesting and re-fusing the same image — a finding about the live stack, not about this recorder.
Related¶
- TASK 20 — the viewpoint-based benchmark this complements (does not replace).
- TASK 21 / 22 — the lift and fusion fixes whose gains should be re-measured on this dataset.
docs/tasks/backlog.md— B2, B3 remain open and are unaffected.