Skip to content

TASK 44 — The ceiling that counted the thinking

A grounding call died with the error the code raises on itself:

RuntimeError: reply hit max_tokens (4096) and is truncated
— raise it rather than treating this as a bad reply

That message is TASK 40's work: truncation used to arrive as "unparseable reply", which reads as a model failure, so ask_claude and ask_gemini both check the stop reason and say plainly that the budget ran out. The check did its job. What it reported, though, is not the failure TASK 40 anticipated.

Not a longer answer — a thinking model against an old ceiling

The number in the parentheses is usage.output_tokens, and it came back equal to the ceiling. A reply cut off late spends most of the budget on the answer and stops mid-object; this one spent all 4096 and produced nothing to cut off. Grounding replies were measured at 460 output tokens on v3, and v4 pushed a five-lantern japanese_room answer past 2048 — nothing about v4 to v6 takes an answer from ~2000 to 4096.

The cause is the model, not the prompt. The default is claude-opus-5, and on Opus 5 thinking is on unless it is switched off — a change from 4.8 and 4.7, where omitting the thinking parameter meant no thinking at all. Thinking is billed against the same max_tokens ceiling as the visible text, so a call that reasons its way across four 100°-wide faces can exhaust the budget before it writes a brace.

This is exactly the shape TASK 40 documented on the other backend, where the 3.x Gemini models think without being asked and the fix was XIAO_HEI_GEMINI_MAX_TOKENS at 8192. The note in the runbook called it one of "two things that differ from Claude". It no longer differs.

What changed

scripts/vlm_probe.py, ask_claude — the only Anthropic call site in scripts/, so execute_plan.py, approach_loop.py, decompose.py and the offline probes all inherit it:

max_tokens=int(os.environ.get("XIAO_HEI_CLAUDE_MAX_TOKENS", "16000")),

16000 rather than something larger: the SDK's non-streaming requests hit an HTTP timeout well below the model's 128000 output ceiling, and going past ~16000 means moving this call to messages.stream(). That is a bigger change than this failure justifies, and 16000 is roughly four times the thinking-plus- answer spend that blew the old ceiling.

The error text now names the variable to raise and says the count includes thinking, matching the Gemini message beside it.

Docs: the runbook gains the variable next to the --model row, a troubleshooting entry keyed to the error text, and a correction to the Gemini section that no longer claims this is a Claude/Gemini difference.

Not verified against the API

No Anthropic key was present in the session that made this change, so the fix is reasoned from the documented Opus 5 default and from output_tokens landing exactly on the ceiling — not from a re-run. The first real drive on this branch is the test. If a call still reports truncation at 16000, the budget is not the whole story and the next step is thinking.display: "summarized" on one call to see what the reasoning is actually spending.

Files

file change
scripts/vlm_probe.py max_tokens 4096 → XIAO_HEI_CLAUDE_MAX_TOKENS (16000); error names the variable
docs/guides/drive-loop-runbook.md variable documented, troubleshooting row, Gemini section corrected