Skip to content

Runaway thinking

Three requests hit exactly 65,536 completion tokens. 240, 252 and 287 seconds each. Thirteen minutes — 12% of a 1.75 hour window — in three requests, none of which delivered an answer.

The signature was consistent. About 249 progress events per run, 100% in the thinking phase, zero answering, zero tool calls. It thinks until the budget dies and never emits anything.

reasoning_chars / completion_tokens ≈ 4.06 on all of them, which means essentially the entire output budget went to thinking.

What actually enables it

reasoning_effort is a sentence in the chat template. It is not a cap. Thinking and the answer draw from the same max_tokens pool, and the client hands over 65,536 by default. The model is politely asked to be brief, and it can still spend all 65,536 tokens being not brief.

Warning

The 65,536 came from my own client. It's Hermes' custom provider-profile floor — default_max_tokens=65536, returned whenever model.max_tokens isn't set. I got this wrong twice before tracing it. It isn't a Strata default and it isn't a foreign client.

The A/B that settled it

Same prompt, same params, only reasoning_effort varied. Hard prompt — tile a 4x10 rectangle with dominoes, derive carefully, check two ways:

sent wall completion reasoning chars visible answer
omitted (server default xhigh) 85.1 s 8192 (cap) 12,831 none
low 71.5 s 8192 (cap) 21,319 none
xhigh 69.1 s 8192 (cap) 17,629 none
none 34.5 s 4,593 0 yes

low was not more restrained than xhigh — it thought longer. On an easy prompt all three levels produced a few hundred reasoning chars with no ordering at all, i.e. noise.

So changing the default level from xhigh to low is a measured no-op. Only none does anything: thinking off, answer delivered, 34 seconds instead of a hang.

Rate, not cause

I spent a while suspecting the 2-bit quant. That was wrong. The same failure shows up on a 3-bit quant on a 12GB laptop, and it happened seven times before this engine existed, on the 27B — five of those seven were 12 to 13.6 minutes each, worse than the 240–287 second runs.

What differs is the rate:

deployment window requests hits ≥60k per 1k wall per hit
27B, llama.cpp Aug 17 – Sep 29 62,776 7 0.11 720–824 s
27B nvfp4, NInfer Sep 22 – 28 18,127 0 0.00 —
125B MoE, Strata ~1 day 408 4 9.80 240–287 s

Neither engine had a thinking budget configured — I checked the launch args, the compose files, and the env. So the ceiling is an engine-agnostic gap. The quant may amplify the tendency; that would need a same-prompt A/B across both engines to prove, and I haven't done it.

What I'd do

  1. Thinking off is the only lever that measurably kills the runaway. Surgical form, keeping thinking on for cloud models and off for the local one: agent.reasoning_overrides: {"qwen3.8-27b": "none"}. The tradeoff is real — agentic turns spend 14k–20k reasoning chars before their tool calls, so this removes chain-of-thought entirely.
  2. Bound the wall clock with model.max_tokens (say 16384). This caps thinking and answer together since they share the pool. It doesn't produce an answer on the hardest prompts; it converts a 250–400 second hang into a 60–100 second failed turn.
  3. Don't bother changing the level. Measured no-op.

A budget buys you an answer. It does not buy a correct one — the domino answer came back confidently wrong with a bogus derivation. Usually the right trade for an agent, since a wrong step gets corrected next turn and a dead turn costs minutes. Wrong trade for a one-shot hard question.

Comments