TASK 26 — Scan-accumulator keyframe sweep and side-by-side dump comparison¶
Why¶
ScanAccumulator keeps the last k registered sweeps and lifts detections
against their union. k had never been justified — it was 10 because it was
10. Tuning it was impractical: every trial meant re-running the sidecar
(detection ~1 s/frame, ~7 min/scene on arabic_room) before anything could be
measured.
Making the sweep cheap¶
Detection depends only on (image, class list, score threshold). None of
those is what a keyframe sweep varies, so the masks can be computed once and
replayed. dump_detections.py already froze them to detections.npz — it did
not need writing, only finding. (A per-run cache was built first and thrown
away; see "Wrong turns".)
replay_score.py and dump_debug.py gained --scan-keyframes,
--scan-voxel and --no-frozen. With frozen masks present, the sidecar is
never started at all. Live vs frozen produced byte-identical metrics, which
is the check that makes the whole shortcut legitimate.
perception_benchmark/sweep_keyframes.sh drives it: freeze once, then loop.
DUMP=1 writes viewer dumps instead of scores.
The sweep¶
DUMP=1 K_VALUES="0 1 2 3 4 5 6 7 8 9 10" sweep_keyframes.sh arabic_room —
11 dumps, 316 frames each, ~97 min, 3.5 GB. Detections are constant at 7482
by construction, so every difference is the accumulator.
| k | final nodes | lifts | lift rate |
|---|---|---|---|
| 0 | 214 | 3829 | 51.2% |
| 1 | 194 | 3585 | 47.9% |
| 2 | 252 | 4149 | 55.5% |
| 3 | 276 | 4471 | 59.8% |
| 4 | 284 | 4681 | 62.6% |
| 5 | 300 | 4841 | 64.7% |
| 6 | 316 | 4956 | 66.2% |
| 7 | 315 | 5081 | 67.9% |
| 8 | 329 | 5175 | 69.2% |
| 9 | 323 | 5246 | 70.1% |
| 10 | 327 | 5311 | 71.0% |
k=0 and k=1 are not adjacent settings. k=0 disables accumulation and lifts the raw sweep (~10,619 points); k=1 accumulates one sweep and voxel-downsamples it to ~5,273. k=1 is therefore the sparsest configuration in the sweep, not the second sparsest, which is why it dips below k=0 on both columns. Read the curve from k=2 up.
From there the trade is monotonic and unhelpful on its own: more keyframes lift
more detections (51% → 71%) and produce more nodes (214 → 327, against 69
GT objects). Denser clouds push marginal detections past min_inliers, and
those weak lifts land as new fragments rather than reinforcing existing nodes.
Node count alone cannot say which end is better — that needs scoring, and it
needs eyes.
Cost per k rose 7:04 (k=0) → 10:18 (k=9): the lift scales with the accumulated cloud, ~5 ms/detection at k=1 against ~76 ms at k=40.
Comparing the dumps¶
viz_app.py gained a Compare runs tab, because the metric table says which
run scores better and not whether the extra nodes are objects or fragments.
PERCEPTION_DEBUG_DIR's siblings are discovered and offered as a sidebar "Dump directory" selector, so a sweep is browsable without restarting.- The compare tab puts N runs side by side at the same viewpoint, sharing the sidebar's viewpoint slider and display toggles. Panels can show the cumulative 3D graph, the per-viewpoint graph, or the overlay image.
- A summary table covers every run (
run_summaryreads scalars only and ispersist="disk"cached — the full table would otherwise hold ~130 MB ofviz.jsonand re-parse it on each app start). - A cumulative-node-count chart per run, against the GT line.
A sidebar class filter restricts phases 3 and 4 — both 3D panels, the inspect picker, the lift-input list, the detection table, and the compare panels — to chosen labels; empty means all. Counts read "1 of 6 nodes" so a filtered view can never be mistaken for the whole graph. Phases 1–2 are unfiltered: the overlay is a PNG the dump already baked. The accumulated point cloud is unfiltered too, and says so — lifted points carry no label once merged.
The viewpoint navigator replaces the bare slider: first/prev/next/last
steppers, a "jump to vp #" box, and ◀ new node / new node ▶, which skip to
the next viewpoint where the cumulative graph gained a node of the selected
classes. That last pair is the point of it — arabic_room is stationary for
two thirds of its 316 frames, so most viewpoints show nothing new and finding
the ones that do by dragging a slider is hopeless. Filtered to potted plant,
54 of 316 frames are growth frames.
The index lives in session_state; every button writes it from an on_click
callback, the only place Streamlit permits assigning a widget-backed key. It is
clamped whenever the run or scene changes, because a shorter capture would make
the slider raise before anything rendered.
One trap in the lift-input picker: its options are the original detection
indices, because sel indexes into the npz point arrays (inmask_det == sel).
Renumbering a filtered list there would select a different detection's points
and quietly show the wrong cloud. When a filter matches nothing at a viewpoint
the list degrades to all classes with a notice rather than calling st.stop(),
which would abandon the rest of the page including the other tabs.
Two details that matter for it being an honest comparison:
Shared axis ranges. With aspectmode="data" each panel scales to its own
extent, so a run with one stray distant node silently shrinks the room and the
panels stop being comparable by eye. Ranges are computed once from the GT
extent and the robot path — identical across runs — and applied to all panels.
Camera presets, not sync. Separate plotly figures cannot share a camera
without custom JS; orbiting one does not move the others. A preset selector
(isometric / top-down / free) puts them all in the same pose instead, with
uirevision keyed on the preset so stepping the viewpoint slider does not
throw away an orbit the user set up.
Node ids are not comparable between panels — they are per-run counters. Labels and positions are.
Animating one class over time¶
perception_benchmark/make_video.py renders the cumulative graph frame by
frame, filtered to a class. The viewer answers "what does the graph look like
at viewpoint N"; this answers "how did it get there", which is the question
fragmentation actually raises — 54 potted plant nodes for 5 real plants were
not found at once.
uv run --with imageio --with imageio-ffmpeg python \
perception_benchmark/make_video.py \
--scene arabic_room --labels "potted plant" --run perception_benchmark/debug_k10
Reads a dump directory only (no sidecar, no GPU); ~37 s for 316 frames.
Output goes to perception_benchmark/videos/ (gitignored).
MP4, not GIF. 316 viewpoints is far too long to watch without a scrubber,
a pause and a speed control, and a GIF has none of them. It is also 10× the
file: the same animation is 7.9 MB as a 159-frame GIF and 0.8 MB as a
316-frame h264. --format gif remains for pasting somewhere that will not take
a video. There is no ffmpeg on the benchmark box, so the encoder comes from the
imageio-ffmpeg wheel, which ships its own binary — hence the --with flags,
matching how the viewer already pulls streamlit and plotly.
viz_app.py lists any clip for the current scene in an Animations expander
and plays MP4s inline with st.video, so the animation sits next to the graph
it animates.
The stream is written frame by frame straight off the matplotlib canvas — no
PNG round-trip, no frame buffer held in memory. h264 needs yuv420p for
anything to play it in a browser, which needs even dimensions; frames are
cropped by a pixel rather than rescaled, so the plot is never resampled.
Nodes that first appeared within the last --new-window viewpoints are drawn
gold, everything older fades to thin crimson. Without that the final frames
are an undifferentiated pile of overlapping wireframes and nothing about the
growth is visible. Node ids are stable in the cumulative map, so first
appearance is just the earliest viewpoint whose nodes list contains the id.
Three rendering details were needed for it to be watchable:
- No
bbox_inches="tight". It crops to drawn content, so the frame size changes as boxes appear — a jittering GIF, and an encoder error for MP4, which needs one constant resolution. set_box_aspectfrom the data ranges. Matplotlib's 3D axes default to a cube, which draws 1 m of room height as long as 9 m of floor — every box reads as a tall column.--z-scale(default 2) exaggerates height on top of the true ratio, because a true-aspect room is nearly flat on screen.- The frame counter says "54 of 327 nodes", not "54". With dozens of
overlapping boxes a filtered frame looks exactly like an unfiltered one —
it was read as "you are drawing every class" on first review, and the
denominator is the only thing that says otherwise. The title is a
fig.suptitle, notax.set_title, because a 3D projection overflows its axes rect and the pane edges draw over an axes-level title (an opaque bbox loses too — it lives inside the same axes). Figure artists draw last. - Fixed limits for the whole animation, computed once from GT extent, the robot path and the final node set. Per-frame limits make the room breathe as nodes appear.
What it shows on arabic_room: plants spawn steadily to ~50 nodes by vp_200
and then plateau, and several of the new nodes are enormous — the largest
has a 5.4 m diagonal, against a real plant's ~0.8 m. Half of the 54 are
oversized boxes (median diagonal 1.46 m, 90th percentile 3.51 m), so the
fragmentation is not only duplicate plants but mask spill fusing a plant with
whatever is behind it. That is a nms_gap/lift problem, not a keyframe one.
Wrong turns¶
A _DetectionCache was written from scratch before checking whether the repo
already froze masks. It did, in dump_detections.py, with two existing
consumers. Deleted.
The first read of the frozen file used allow_pickle=False and crashed:
labels is dtype=object. Both existing consumers already passed
allow_pickle=True. Caught by a 4-frame probe run before committing to a
40-minute sweep — worth the two minutes.
Verification¶
streamlit.testing.v1.AppTest drives the app headless: it renders with no
exceptions, exposes all three tabs, defaults the comparison to the selected run
plus its sweep neighbour, and survives stepping the viewpoint slider (which
re-renders every compare panel). Lint is unchanged from HEAD at 10.
Open¶
The scored sweep has not been run — only the dumps. sweep_keyframes.sh
arabic_room without DUMP=1 reuses the same frozen masks and would score all
11 in ~35 min. Until then the qualitative review is the only evidence about
which k to pick.