Skip to content

The number that lies to you: decode tokens per second

A 125B MoE reported better decode than the 27B it replaced, and it felt slower. Both things were true. The engine log wasn't lying — it was answering a different question.

What the log reports

The tok/s in the engine log is decode-only, sampled mid-request. It doesn't include the time spent reading the prompt. Over 270 metered requests:

Strata 125B NInfer 27B
decode-only tok/s, p50 165 154
effective tok/s (completion / total duration), p50 119 130
overhead factor 1.39x 1.19x

The 125B is about 7% faster on decode and about 8% worse end to end. The overhead is prefill, and prefill is where the cache lives or dies.

The tail is the whole story

Medians are close. The tail isn't:

NInfer 27B Strata
ttft p50 / p99 / max 281 ms / 16.2 s / 66.9 s 688 ms / 51.2 s / 137.9 s
total duration p50 / p90 / p99 / max 3.1 s / 18.0 s / 54.1 s / 148.5 s 3.8 s / 23.7 s / 240.0 s / 286.8 s
reasoning chars p99 22,414 266,325

p50 says the two are a wash. p99 says one of them is four to five times worse. When you say "it feels slower," you're not perceiving the median — you're perceiving the tail. Nobody experiences a p50.

Where I got this from

Not the engine log. The proxy's metering DB — one row per request with prompt_tokens, cached_tokens, prompt_ms, predicted_ms, duration_ms, ttft_ms, ttft_visible_ms, reasoning_chars, finish_reason, served_model.

Tip

ttft_visible_ms being NULL is its own signal: it means the turn never delivered user-visible content. Every token went to reasoning_content. That column alone tells you whether a request produced an answer.

If you want to separate cache damage from generation cost, decompose before blaming either:

pref = sum(r['prompt_ms'] for r in rows) / 60000      # cache-miss cost, minutes
pred = sum(r['predicted_ms'] for r in rows) / 60000   # generation cost, minutes

Over 269 requests and 105 minutes of wall clock: 48.9 minutes of summed request time, 9.5 minutes prefill (19.5%), 33.1 minutes generation (67.7%). Two independent problems, roughly 2:1 in favour of thinking.

The proof they're independent: one runaway ran with cached=167,664 — a 99.5% hit — 0.5 seconds of prefill, no competing agent, and still burned 65,536 tokens over 252 seconds. A clean cache does not prevent a runaway.

What I'd benchmark instead

Report p99, not p50, and report effective throughput rather than decode throughput. And measure with two clients, because a single-client benchmark can't show eviction at all — there's nothing to evict.

Comments