The Checkpoint Sidecar Patch¶
This one is worth writing down because it's a bug you won't find in upstream, and it cost me a few build cycles to pin down.
The symptom¶
After a slot restore, the first request comes back at 0% cached. The second request is warm at 99%+. Everything looks healthy, the KV rehydrates perfectly, and then the first turn pays a full re-prefill anyway.
Root cause¶
slot.prompt.checkpoints live only in process memory. llama_state_seq_save_file never serializes them. So after SLOT_RESTORE the checkpoint list is empty, and the first request's rollback — the BPE re-tokenize of the prompt tail — finds no covering checkpoint:
The KV is there. The map telling the server which KV pages are valid for the prompt tail is not.
The fix¶
PR #206 on TheTom/llama-cpp-turboquant (apollo-mg, commit eaf98e612) persists checkpoints to a <state>.ckpt sidecar file (magic PKCL) at SLOT_SAVE and reloads them at SLOT_RESTORE. About 117 lines in one file.
It is not in upstream ggml-org master — I verified that on 2026-08-24. Bumping your pin will not pick it up. Tracking upstream: ggml-org PR #24028.
Results¶
A/B on a CPU rig, single slot, lfm2-5-1-2b: first-after-restore went from 0.00% to 99.48%. The patched log says it plainly:
On the prod box (llama-server-sidecar:2026.08.24-v4) the live test landed at 72–74% cached first-after-restore, which is where I'd expect it to sit with real agent traffic.
Gotchas applying it¶
Two things that bit me:
- The raw fork patch doesn't apply cleanly to a pin. Regenerate it with
git apply --3waythengit diff --cached. - There's a semantic gap: the fork's parent has a
token_countlocal that doesn't exist in my pin. Replace it with!slot->prompt.tokens.empty()or it won't compile.
Also, a separate issue still open: stale q4_0 .bin profile files left over from a cache-type flip fail restore with mismatched key type (8 != 2). Re-save all profile files after any cache-type change.
The image is locally built¶
The patched image is a locally built image, not the stock ghcr.io/ggml-org/llama.cpp:server-cuda. That matters because a compose file pointing at the stock image silently reverts the fix — first-after-restore cache goes back to 0% and nothing looks wrong. I hit exactly that when a branch pointed at the stock image while prod ran the patched one. Pin the sidecar tag, and check that the compose file matches what prod actually runs.