TASK 28 — Swap YOLO-World for OWLv2 (large) and drop the tick to 1 Hz¶
Outcome: OWLv2 reverted; back to YOLO-World. The on-GPU A/B on
arabic_room(397-framecaptures_nav, same frozen captures, GT = 69 objects) did not justify the swap:
detector @ thresh mAP@1m P R F1@1m cErr (m) pred count YOLO-World @ 0.25 0.397 0.443 0.449 0.446 0.160 70 ≈ GT OWLv2-large @ 0.1 0.649 0.191 0.768 0.306 0.198 277 (4× over) OWLv2 wins ranking (mAP) and recall but massively over-produces at the 0.1 floor — precision 0.19, F1 and counting both worse, and the predicted node count (277) is ~4× the ground truth (69), which is what made the maps "look worse". On top of that OWLv2-large ran ~4.6–6 s/frame on the A10G, which does not fit even the relaxed 1 Hz tick. The high mAP says the ranking is good, so threshold tuning might have recovered precision — but not the latency — so we reverted the detector and kept the two orthogonal wins.
Kept from this task (not reverted): the 1 Hz tick (a good change on its own), the SAM 2.1 Hiera Large upgrade (detector-agnostic, closes the old B5 SAM-Tiny gap), the
sam_score(SAM predicted-IoU) plumbing, and theviz_app.py/sweep_gates.sh/dump_debug.pybenchmark tooling. Reverted: the box detector (OWLv2 → YOLO-World),transformers→ultralytics, the Dockerfile's OWLv2 bake, and the score-threshold floor (0.1 → 0.25) across the request path and benchmark scripts.
Why¶
Two coupled changes to raise perception quality:
- Detector. YOLO-World was chosen for speed, which stops mattering once Task 1/2 answer within a 10-minute budget rather than in real time. Swapping it for a heavier open-vocab detector trades latency (which we have) for recall on the challenge's object vocabulary (which we need). We keep the box→SAM architecture and only replace the box stage — this is the Grounded-SAM-2 pattern.
- Tick rate. OWLv2 + SAM-Large no longer fits the old 500 ms / 2 Hz tick, so the default drops to 1 Hz (1000 ms). The lower rate is what buys the latency budget for the heavier detector; the two changes reinforce.
What changed¶
Detector: OWLv2 large (google/owlv2-large-patch14-ensemble)¶
perception/pipeline.py- Load
Owlv2Processor+Owlv2ForObjectDetectionin place ofultralytics.YOLOWorld(lazy import inside__init__, as before). set_classesnow just caches the class list — OWLv2 has no persistent prompt to pre-encode; queries are tokenised + encoded jointly with the image on everydetectforward (cheap for a handful of classes at 1 Hz).detectruns OWLv2 on the 4-face batch, post-processes withpost_process_object_detection(boxes come back in each face's own pixel coords), then applies per-face NMS (torchvision.ops.nms) — OWLv2 has no built-in NMS, unlike YOLO-World'siou=arg. Boxes are clamped toFACE_SIZE. The SAM-per-box path and seam merge are unchanged.- Faces are square (
FACE_SIZE), so OWLv2's square-pad is a no-op and the known non-square box-offset caveat does not apply. - The box→SAM interface,
sam_score(B5), range cap (B4), and seam merge are all detector-agnostic and unchanged.
Score threshold rescaled to OWLv2's range¶
OWLv2's query-match scores run lower than YOLO-World's box confidence, so the
detection floor moved 0.25 → 0.1 consistently across the request path:
responder.DEFAULT_SCORE_THRESHOLD, server.py form default, client.detect,
pipeline.detect, and the two benchmark scripts
(replay_score.py, dump_detections.py). This is a starting point — tune on
the frozen captures with replay_score.py --score-threshold.
Sidecar image¶
perception/requirements.txt:ultralytics→transformers==4.46.3.perception/Dockerfile:- pip install
transformers; dropped the OpenAI-CLIP install (it was YOLO-World's text encoder) and theCOPY ViT-B-32.ptstep. - Bake SAM 2.1 Hiera Large (was Tiny — this also closes the B5 gap where
the code referenced
sam2.1_hiera_large.ptbut the image only fetched Tiny) and OWLv2 weights into the HF cache. HF_HUB_OFFLINE=1 / TRANSFORMERS_OFFLINE=1set after the bake, so a live tick never blocks on the network during the challenge.
Tick rate: 2 Hz → 1 Hz¶
src/xiao_hei_vln/app/main.py:XIAO_HEI_VLM_TICK_HZdefault2.0→1.0(timer period is1.0 / TICK_HZ). Override still honoured.perception_benchmark/record_navigation.py+run_nav_capture.sh: mirror the default so recorded captures match the live cadence.- Docs:
README.md,docs/getting-started/configuration.md,docs/concepts/tick-loop.mdupdated to the 1 Hz default (incl. the "Why 1 Hz?" budget table).
Status¶
Measured on-GPU, then reverted — see the Outcome banner at the top. OWLv2
was swapped in, the image rebuilt, and the arabic_room A/B run: OWLv2's 4×
over-production at the 0.1 floor (277 vs 69 GT) and ~5 s/frame latency lost to
YOLO-World on precision, F1, counting, and tick fit despite a higher mAP. The
detector was reverted to YOLO-World; the 1 Hz tick and SAM-Large upgrade were
kept. Sidecar image rebuilt on the YOLO+SAM-Large definition; all unit tests
pass (perception responder/client/sync + sidecar geometry/seam-merge:
76 + 51 = 127).
Next¶
- If OWLv2 is revisited, first tune the score threshold (the 0.649 mAP says ranking is good) and quantise / use a smaller OWLv2 variant to get under the tick budget before re-running the A/B.
- With YOLO-World restored, layer the B4/B5 gate sweep (
sweep_gates.sh) — now backed by realsam_scorevalues from SAM-Large — on the frozencaptures_navdumps. - Re-baseline
arabic_roomwith YOLO + SAM-Large (vs the old YOLO + SAM-Tiny) to isolate the mask-quality gain the upgrade should bring.
Related¶
- TASK 27 — time-sync + extrinsic fix (the lift this detector feeds).
- Backlog B4 (range cap) / B5 (SAM mask quality + SAM-Large) — the gates this composes with.