Why two agents on one local LLM feel slow¶
I ran this for a while with two agents on the same engine and both of them felt sluggish, and I couldn't work out why. The engine reported good numbers. The experience didn't match.
The answer is the prefix cache, and it's structural — not a tuning knob.
One window, no second tier¶
Strata's prefix cache is a single 262,144-token KV window in VRAM. About 3.09 GiB at int8. There is no host-RAM tier behind it.
llama.cpp has --cache-ram, which gives it a host-RAM checkpoint tier. That single difference is the whole story:
| cache hit | cache miss | |
|---|---|---|
| Strata | prefill p50 408 ms | 10.4% of requests, prefill p50 15.0 s (p90 29.6 s, max 31.2 s) |
| llama.cpp / NInfer | prefill p50 265 ms | 5.4% of requests, prefill p50 6.5 s |
Now put two agents on it. One at 164k tokens, the other at 190k. That's a 354k working set into a 262k window. Mutual eviction is guaranteed. Not likely — guaranteed.
The proof was in the interleaving: 7 of the last 8 misses came immediately after a request from the other key. Same key back to back, cached_ratio 1.00 every time. Worst stretch I saw was one agent missing six consecutive turns at 147k–190k prompts, paying 23–31 seconds of prefill each. Nearly three minutes of prefill inside five minutes of wall clock.
Watch for cached_tokens = 16,384. That's exactly one chunk. It reads like a partial win and it's a near-total miss.
What fixed it¶
Upstream landed a conversation cache in v0.1.30 (tagged 2026-09-30) that parks whole snapshots in host RAM instead of VRAM:
Two alternating conversations, 57k and 60k tokens, reasoning_effort: none, temp 0:
| turn | prompt | cached | prefill | wall |
|---|---|---|---|---|
| A1 cold | 57,122 | 0 | 9,170 ms | 9.3 s |
| B1 cold | 59,721 | 0 | 10,034 ms | 10.1 s |
| A2 — the turn that used to be a guaranteed re-read | 57,124 | 49,152 | 2,098 ms | 2.2 s |
| B2 | 59,723 | 49,152 | 2,259 ms | 2.4 s |
evictions=0 with both conversations live. Prefill on re-entry dropped about 4.4x, and the restore itself costs 75 ms.
Two honest caveats. Reuse resumes from a checkpoint — taken every 16K tokens and at each assistant turn — so the tail after the boundary gets re-read. That's where the 2 seconds come from. And issue #342 is still open upstream: successive turns of the same conversation each consume a slot, so a long subagent chain can push the parent out. Alternating different chats works today. Delegation chains don't.
Sizing it¶
Measured: about 1.5 GB per 57k-token conversation. That's roughly 26 KB per token, about double the int8 KV's 12.7 KB, because a snapshot carries running state, checkpoints, GDN/SSM and indexer state, and draft KV — not just KV.
RAM is the binding constraint here, not VRAM. With 48 GB total and about 8 GB available while serving, the 8 GiB example budget in the docs gets refused by the 2560 MiB floor. Expect 4–5 GiB, which is roughly two conversations at that size. Two agent chats at 150–190k tokens need 8–12 GB.
What I'd do with this¶
If you're running one agent with long context, this is fine and the numbers hold. If you're running two agents that both exceed ~130k tokens, no flag fixes it — route one of them to an engine with a host-RAM tier, or size the conversation cache to actually fit both.