The number that lies to you: decode tokens per second¶
A 125B MoE reported better decode than the 27B it replaced, and it felt slower. Both things were true. The engine log wasn't lying — it was answering a different question.
What the log reports¶
The tok/s in the engine log is decode-only, sampled mid-request. It doesn't include the time spent reading the prompt. Over 270 metered requests:
| Strata 125B | NInfer 27B | |
|---|---|---|
| decode-only tok/s, p50 | 165 | 154 |
| effective tok/s (completion / total duration), p50 | 119 | 130 |
| overhead factor | 1.39x | 1.19x |
The 125B is about 7% faster on decode and about 8% worse end to end. The overhead is prefill, and prefill is where the cache lives or dies.
The tail is the whole story¶
Medians are close. The tail isn't:
| NInfer 27B | Strata | |
|---|---|---|
| ttft p50 / p99 / max | 281 ms / 16.2 s / 66.9 s | 688 ms / 51.2 s / 137.9 s |
| total duration p50 / p90 / p99 / max | 3.1 s / 18.0 s / 54.1 s / 148.5 s | 3.8 s / 23.7 s / 240.0 s / 286.8 s |
| reasoning chars p99 | 22,414 | 266,325 |
p50 says the two are a wash. p99 says one of them is four to five times worse. When you say "it feels slower," you're not perceiving the median — you're perceiving the tail. Nobody experiences a p50.
Where I got this from¶
Not the engine log. The proxy's metering DB — one row per request with prompt_tokens, cached_tokens, prompt_ms, predicted_ms, duration_ms, ttft_ms, ttft_visible_ms, reasoning_chars, finish_reason, served_model.
Tip
ttft_visible_ms being NULL is its own signal: it means the turn never delivered user-visible content. Every token went to reasoning_content. That column alone tells you whether a request produced an answer.
If you want to separate cache damage from generation cost, decompose before blaming either:
pref = sum(r['prompt_ms'] for r in rows) / 60000 # cache-miss cost, minutes
pred = sum(r['predicted_ms'] for r in rows) / 60000 # generation cost, minutes
Over 269 requests and 105 minutes of wall clock: 48.9 minutes of summed request time, 9.5 minutes prefill (19.5%), 33.1 minutes generation (67.7%). Two independent problems, roughly 2:1 in favour of thinking.
The proof they're independent: one runaway ran with cached=167,664 — a 99.5% hit — 0.5 seconds of prefill, no competing agent, and still burned 65,536 tokens over 252 seconds. A clean cache does not prevent a runaway.
What I'd benchmark instead¶
Report p99, not p50, and report effective throughput rather than decode throughput. And measure with two clients, because a single-client benchmark can't show eviction at all — there's nothing to evict.