TASK 25 — A viewer for the perception map, and what it found¶
Every number in TASK 22–24 is an average over hundreds of observations, and an
average cannot show why a box is wrong. Twice already a picture has overturned
a conclusion the aggregates supported — the equirect overlay is what caught the
fan decoration ground truth being one box around three fans plus the wall
between them, which is what made TASK 23's 1.622 lidar oracle partly circular.
So this task built a browser viewer for the replayed scene, and then used it. Four measurements came out of the exercise. Three of them contradict things this repository currently believes, including one that says our reported counting error has been inflated by a bug in the metric rather than by the pipeline.
The viewer¶
scripts/export_viz.py replays a scene through the real PointLifter and
ObjectMap at the benchmark's gating (min_move 0.15, min_rot 10°,
min_score 0.35, min_inliers 10) and dumps a JSON manifest plus one
float32 blob of xyz triples. The manifest indexes into the blob by point
offset, so the page fetches geometry as an ArrayBuffer and never parses
coordinates out of text. A box on screen is the same box in the benchmark
table, not a re-derivation free to drift from it.
scripts/serve_viz.py serves viz/ and fetches three.js r128 into
viz/vendor/ on first run. r128 is the last release whose builds are plain
<script> globals rather than ES modules, which keeps the page dependency-free.
Layers: accumulated lidar cloud, the current frame's scan, the lifted inlier points per detection (coloured per label), per-frame detection boxes, our fused objects, ground-truth boxes, the ground-truth scene cloud, and the robot path. Frame slider with play, filters on label / score / observation count, and click-to-inspect on any box.
Two things worth knowing about how it works:
- Frame provenance.
ObjectMap.addpicks or creates a node without telling the caller which, so the frame stamp is taken from the node side by wrapping_Node.__init__and_Node.merge. That covers both the born-here and merged-into paths without duplicating the merge rule, which would be free to drift from the real one.node_idsurvivesfinalize(), so the stamps still key the exported objects. This drives the detected up to this frame and seen in this frame modes — the only way to watch a duplicate node split off as it happens. - The ground-truth cloud is each scene's own
map.ply, read straight out of the challenge zip and voxel-thinned to 5 cm. It is barexyz: no colour, no per-point labels.loftis 12.7 M points thinned to 443 k, and the whole seven-scene export costs ~20 MB extra.
Everything runs locally: rsync the ~130 MB of viz/data/ to any machine and
serve_viz.py needs nothing else. viz/data/ and viz/vendor/ are gitignored.
Finding 1 — our counting error is inflated by the metric, not the pipeline¶
cMAE averages |gt_c - pred_c| over every class either side names. That one
number hides four different failures, and only two are duplicates:
| term | meaning | who owns it |
|---|---|---|
over |
class exists, we predict too many | deduplication |
spurious |
class ground truth does not have | naming |
short |
class named, too few found | recall |
absent |
class never named at all | vocabulary |
Decomposed over the seven benchmark scenes (per-scene means):
| config | #pred | over | spurious | short | absent | cMAE | floor |
|---|---|---|---|---|---|---|---|
| greedy 0.4 (shipped) | 441 | 99 | 299 | 3 | 32 | 7.27 | 0.59 |
| merge_dist=0.8 | 262 | 48 | 174 | 6 | 32 | 4.44 | 0.65 |
| relative=1.0 | 254 | 63 | 150 | 4 | 32 | 4.22 | 0.61 |
| families | 415 | 87 | 289 | 3 | 35 | 7.09 | 0.65 |
| fam+0.8 | 237 | 39 | 163 | 5 | 37 | 4.28 | 0.74 |
| fam+0.8+n_obs>=3 | 174 | 23 | 119 | 6 | 40 | 3.77 | 0.91 |
| fam+recl0.8+n_obs>=3 | 154 | 19 | 104 | 6 | 40 | 3.42 | 0.93 |
floor is (short + absent) / classes: what survives if every duplicate and
every spurious node vanished at no cost to what we did find. It is 0.59, not
the ~4 that had been assumed. Recall is not what holds cMAE up.
But the dominant term, spurious, was mostly an artefact of the measurement.
The ground truth is filtered by is_structure; the predictions were not. Score
both sides the same way:
| variant | #pred | over | spurious | short | absent | cMAE | floor |
|---|---|---|---|---|---|---|---|
| as scored today | 441 | 99 | 299 | 3 | 32 | 7.27 | 0.59 |
| structure dropped from predictions too | 248 | 99 | 107 | 3 | 32 | 4.59 | 0.67 |
Of 2095 surplus nodes across seven scenes, 1348 (64%) are structure —
floor×537, ceiling×428, window×100, carpet×90, door×75, wall×72,
column×27 — and 1309 of them carry a label the ground-truth zip does
contain, stripped only by is_structure. These are not detection errors.
Correcting the asymmetry is a metric fix, not an improvement, and it is the
first thing to do.
After the fix, the two movable terms are the same size: over 99 against
spurious 107. Deduplication and naming are equal partners, not a hierarchy.
Finding 2 — the spurious labels are synonyms, not hallucinations¶
The non-structure surplus, by label, across seven scenes:
lamp 71 · dining table 62 · picture 54 · couch 49 · bench 44 · tv 42
side table 31 · bed 25 · cabinet 25 · wardrobe 24 · tv stand 23 · lantern 23
photo 22 · bookcase 22 · painting 19 · coffee table 18 · wall lamp 16
This is almost exactly the FAMILIES table already sitting in
scripts/fusion_sweep.py: couch/sofa, dining table/table, picture/painting/
photo, bench/chair, lamp/lantern/wall lamp/ceiling lamp. The viewer shows it
directly — a ceiling full of overlapping lamp / wall lamp / focus light /
lantern boxes on the same fixtures.
Duplicates inside classes ground truth does have (693 total) are far more
concentrated: chair 192 (28%), potted plant 54, cabinet 46, painting 34.
Finding 3 — a global merge threshold cannot fix counting¶
The class each scene's official numerical question counts, against what each fusion config produces:
| scene | target | GT | shipped | merge 0.8 | rel 1.0 | fam+0.8+n≥3 | fam+recl |
|---|---|---|---|---|---|---|---|
| arabic_room | sofa | 3 | 18 | 11 | 14 | 9 | 9 |
| chinese_room | chair | 6 | 32 | 15 | 17 | 12 | 10 |
| japanese_room | calligraphy painting | 3 | 0 | 0 | 0 | 0 | 0 |
| livingroom_3 | photo | 2 | 2 | 2 | 2 | 1 | 1 |
| loft | pillow | 11 | 8 | 6 | 8 | 6 | 5 |
| office_1 | computer monitor | 6 | 14 | 7 | 9 | 8 | 7 |
| office_2 | potted plant | 3 | 5 | 4 | 5 | 3 | 2 |
Aggressive merging fixes the over-counted classes (sofa, chair,
computer monitor) and breaks the ones already under-counted — photo
2 → 1, pillow 8 → 5 against a ground truth of 11. One global threshold cannot
serve both. This is direct evidence for a selective rule (the size-prior veto
sketched in TASK 24's notes) over a looser uniform one.
japanese_room's calligraphy painting is 0 at every setting. No fusion rule
reaches a vocabulary failure.
Note also that all fifteen official numerical questions are conditional counts — "how many X are on Y" — so the whole-scene count is necessary and nowhere near sufficient.
Finding 4 — a detection box with no object under it, attributed¶
japanese_room frame 0, 17 lifted detections, 244 nodes built, 185 exported:
| verdict | n | meaning |
|---|---|---|
| kept | 9 (53%) | fused object overlaps the detection |
| moved | 3 (18%) | node survived, but its fused box sits elsewhere |
| structure | 3 (18%) | exists; hidden by the structure filter |
| nms | 2 (12%) | deleted outright by finalize() at IoU 0.84 and 0.51 |
The moved cases are the interesting ones: potted plant node 7 (n_obs 22)
has IoU 0.05 with its own frame-0 observation, bird decoration node 6
(n_obs 27) has 0.04. The fused box is the point-count-weighted mean over 20+
views, so it need not sit on any single one — this is what TASK 24's 0.236 m
lateral centre error looks like on screen.
bird decoration also appears twice in that one frame, both merging into the
same node — TASK 24's "30.7% of object-frames have more than one detection".
Finding 5 — pricing the ground-truth box convention¶
object_list.txt annotates an oriented box: id cx cy cz lx ly lz heading
"label". perception/eval.py:286 builds its AABB as center ± size/2 with
the heading ignored. For a rotated object the box we are scored against is
the object's local box dropped into the world unrotated — neither the true
oriented box nor its world-aligned bound.
Our predictions come from point clouds and are world-axis-aligned by construction, so the best a perfect system could produce is the world AABB of the true oriented box. Over 542 non-structure ground-truth objects:
| heading under 5° (unaffected) | 50.7% |
| mean IoU a perfect axis-aligned prediction could reach | 0.783 |
| p25 / p50 / p75 | 0.590 / 0.843 / 0.986 |
| below 0.75 | 42.8% |
| below 0.50 — can never earn the 2-point tier | 6.5% |
Per scene the mean runs 0.735 (chinese_room) to 0.850 (office_2).
This is real but not the current bottleneck: it caps mean IoU at 0.783 and we are at 0.177. It matters later. The exception is the 6.5% that can never reach IoU 0.5 however good the perception gets.
The worst-affected are thin elongated objects — pen 0.139, marker 0.146,
chopsticks 0.247 — which overlaps heavily with the sub-0.3 m bin TASK 23
showed cannot reach IoU 0.5 even when cheating. The penalty largely stacks onto
existing failures rather than adding new ones. The exception worth noting is
loft's pillows: four of the fifteen worst, and loft's numerical question is
"How many black pillows are on the sofa?"
The viewer can draw either convention ("…drawn with heading (true OBB)"), so the difference is inspectable per object.
Corrections to earlier claims¶
- "441 nodes against 542 ground-truth objects, so we under-count." Wrong:
fusion_sweep.pyreportsnas a per-scene mean, not a total. Roughly 3000 nodes against 542 objects — we over-count about 5×. - "Deduplication bottoms out around cMAE 4.0." Wrong: the floor is 0.59
(0.67 scored symmetrically), and
fam+recl0.8+n_obs>=3already reaches 3.42. - "Spurious classes dominate." True as measured, but 64% of that was our own asymmetric structure filtering, not the detector.
The viewer reproduced the same asymmetry on screen at first — per-frame
detections were judged by a narrow regex while fused objects used
vocab.is_structure, so windows and doors drew a detection box with no object
under it. The manifest now carries one structure verdict per label for both.
Next¶
- Make the structure filter symmetric in the benchmark. A metric fix worth
cMAE 7.27 → 4.59 with no pipeline change. Decide at the same time whether to
admit openings on both sides:
window100 +door75 +column27 +window frame12 = 214 nodes per seven scenes would become legitimate, and 22 of the 75 official questions name one. - The alias merge from TASK 23, still unimplemented, against
spurious107. - Selective deduplication — size-prior veto plus relative radius — against
over99, with Finding 3 as the acceptance test: it must not costphotoorpillow. - Vocabulary for the zero-detection classes such as
calligraphy painting.
Files¶
| file | what |
|---|---|
scripts/export_viz.py |
replay one scene into manifest + float32 blob |
scripts/serve_viz.py |
static server, fetches three.js on first run |
viz/index.html, viz/app.js |
the viewer |
.gitignore |
ignore viz/data/, viz/vendor/ |
The four analyses above were run from scratch scripts on the box
(scratch_cmae.py, scratch_cmae2.py, scratch_frame0.py,
scratch_heading.py), uncommitted, in the same style as TASK 23's and 24's.
441 tests pass, 1 skipped; no library code changed in this task.