TASK 50 — Why both explorers stall, and which one to keep¶
Two sweeps were on disk with no verdict attached: exploration_logs_frontier/
(19 scenes, no wall-clock budget) and exploration_logs_nbv/ (13 scenes, the
480 s budget from TASK 49). The question was which strategy to invest in, and
then to fix what the logs showed was broken.
Both have since been renamed to exploration_logs_<strategy>_exp0/ — they are
the pre-fix baseline, and every measurement below comes from them.
Verdict: frontier¶
The budgets differ, so absolute totals are not comparable. Normalising by wall clock on the 12 scenes both strategies ran:
| frontier | nbv | |
|---|---|---|
| mean duration | 353 s | 389 s |
| mean path travelled | 87.8 m | 35.5 m |
| metres per minute | 9.17 | 5.01 |
| mean pose-bbox area | 135.8 m² | 43.7 m² |
| mean waypoints advanced | 16.9 | 4.1 |
Frontier wins metres-per-minute on 9 of 12 scenes. It covers ~1.8× more ground per unit time and did so while terminating early on most scenes — it had no time budget, so its skip-hatch fired at 130–250 s while nbv was still running its full 480 s.
The decisive number is not the ratio, it is why each one fails. Splitting
skips by whether /way_point_reached ever published during the attempt:
| frontier | nbv | |
|---|---|---|
skips where nav stack stayed silent (best_nav_dist=inf) |
409 / 611 | 248 / 487 |
| median robot displacement during those skips | 2.39 m | 0.00 m |
Frontier's failures are a robot that was driving and got interrupted. Nbv's are a robot that could not move at all. The first is a bad supervisor heuristic, fixable in the tick loop; the second is the strategy poisoning its own map. Frontier is the one to keep, and nbv is worth keeping alive only as a comparison baseline.
Root causes and fixes¶
1. A silent nav stack read as "the robot has stopped" (both; biggest win)¶
main.py's early-skip fires when best_nav_dist fails to improve for 5 ticks.
best_nav_dist initialises to inf, and inf >= inf - 0.02 is true, so
whenever /way_point_reached published nothing the counter ran to 5 unopposed
and the waypoint was dropped ~6 s in. That is 97% of frontier's skips and 80%
of nbv's — they are almost all the early skip, not the 12 s stuck timeout.
Fix: track distance-to-target in odometry alongside the nav topic
(best_odom / last_progress_time) and refuse to early-skip while the robot
is visibly closing on the goal. Replaying the recorded skips against the new
predicate: 71% of frontier's early skips and 36% of nbv's were still closing
on their target and would now be deferred to the stuck timeout.
best_odom_dist is added to the WP_SKIP log line so the next sweep is
diagnosable without this reconstruction.
2. Explorers stamped their own waypoints as obstacles (both; nbv-fatal)¶
Both strategies called OccupancyGrid.mark_occupied() on visited and skipped
waypoints to stop re-selecting them. That does not just suppress a goal — it
deletes the cells from _free and blacklists them against future terrain
updates. A waypoint in a doorway becomes a wall. Nbv picks targets through
reachable_path_costs(), a 4-connected BFS over free cells, so every visited
waypoint chopped a 1.4 m hole in its own reachability graph until nothing was
reachable. Hence the 0.00 m median displacement.
Fix: new mark_no_target() / is_targetable() on the grid, backed by a
separate _no_target set. Cells stay FREE and stay traversable; they are only
excluded from selection. mark_occupied() keeps its old meaning and is now
reserved for genuine obstacles.
3. Frontier re-issued one target 18 times in a row (loft)¶
Suppressing the cells under a rejected centroid leaves the rest of the cluster
averaging to nearly the same point, so (3.62,0.62) → (3.64,0.61) →
(3.64,0.61) … until the skip hatch closed the run at visited=0. Frontier
repeated a target 8.7 times per scene on average.
Fix: an explicit _rejected list of failed centroids in world coordinates,
with a reject_radius (0.8 m) test applied at selection time — independent of
what the grid does with the underlying cells.
4. One barren tick ended the sweep as no_frontiers¶
7 of 19 frontier scenes died this way at 130–250 s. A single tick where every
cluster fell under min_frontier_size set _done. Terrain arrives in bursts;
that is noise, not coverage.
Fix: empty_ticks_before_done (20 ticks) before declaring done, and a fallback
to under-sized clusters — a 3-cell frontier is still unseen space.
5. max_consecutive_skips fired while the map was still growing¶
The hatch is meant to catch a stalled sweep, but it counted skips with no reference to whether exploration was still making progress.
Fix: on hitting the limit, compare the free-cell count against the count at the
last advance. If the map has grown by skip_reset_free_cells (60 cells ≈
2.4 m²) the sweep is still learning — bank the progress, reset the counter, and
continue. Otherwise terminate as before.
6. Nbv detail fixes¶
- Skip blacklisted a single 0.2 m cell, so the next pick landed 20 cm away and failed identically. Now suppresses a disc.
_pickdrew 40 uniform samples from a pool where frontier cells were merely weighted ×3. Frontier cells are the only candidates carrying information gain, so they are now all scored nearest-first up ton_samples, with random free-space draws only filling the remainder.
Instrumentation for the next sweep¶
Every number in this report is a proxy. Coverage was reconstructed from pose
bounding boxes, path length from poses sampled only at WP_SET/WP_SKIP, and
"was this the early skip or the stuck timeout?" from an elapsed < 11 s
heuristic. The next sweep should not need any of that:
| Added | Answers |
|---|---|
MAP heartbeat every 10 s (XIAO_HEI_EXPLORATION_MAP_LOG_S) — free, occupied, frontier, frontier_open, no_target, free_m2, reachable, path_m |
The coverage curve, and the gaps between waypoints. reachable falling away from free is a strategy walling itself in — the nbv failure, visible directly instead of inferred |
kind=no_progress\|stuck_timeout on WP_SKIP |
Which of the two skip paths fired, instead of guessing from elapsed |
best_odom_dist on WP_SKIP |
Whether the robot was closing on the target when it was abandoned |
path_m accumulated every tick, on MAP/WP_SKIP/DONE |
True path length; the old pose sampling undercounts worst on the scenes that drive most |
NO_TARGET (throttled to one line per reason-change or 5 s) with why= |
Why the selector came back empty — all_suppressed (over-blacklisted) vs no_frontier_cells (real coverage) vs unreachable demand opposite responses and were indistinguishable |
why= + candidate counts on WP_SET |
What the selector was choosing between |
hatch_resets, select_why, and the final grid counters on DONE |
A one-line summary per run |
Supporting changes: scripts/exploration_sweep.sh collates the new fields into
results_<strategy>.csv and now measures duration from CLOCK_START rather
than START, so the 90-190 s wait for /state_estimation is not billed as
exploration time. scripts/compare_exploration.py prints the per-strategy and
paired per-scene tables above from any log directory — it reads the old logs
too, marking reconstructed columns with ~, and excludes runs that produced no
data (frontier/chinese_room is a START line and nothing else) so a crashed
container cannot lose a head-to-head.
Tests¶
tests/test_exploration_targeting.py — 20 tests. One per failure above:
suppression leaves a corridor traversable while mark_occupied still blocks
it; a visited nbv waypoint does not disconnect the map behind it; neither
strategy re-proposes a skipped target; a barren tick does not end a sweep; the
skip hatch resets while free cells are still arriving. Then one per diagnostic
field, since a log line that lies is worse than no log line: stats()
separates a real obstacle (reachable halves) from target suppression
(reachable unchanged), and each why= value is pinned to the state it
names.
Full suite: 521 passed. (tests/test_execute_plan.py does not import in this
environment — scripts/approach_loop.py needs cv2, pre-existing.)
Not done¶
Nothing was re-run against the simulator; every number above is reconstructed from the two logged sweeps. The next sweep should run both strategies under the same 480 s budget — frontier never had one, which is the single largest confound in the table at the top. Then:
Three things to check in the result, in order:
kind=on the skips.no_progressshould have collapsed as a share of the total. If it has not, the odometry guard is not doing its job and_WP_ODOM_STALL_Sis the knob.reachablevsfreein theMAPlines, for nbv especially. They should now track each other for the whole run. Divergence means something else is still eating the reachable set.free_m2per minute, which is the first honest coverage number either strategy will have produced. If it disagrees with metres-per-minute about which strategy is better, believefree_m2— path length rewards a robot that drives in circles.