Skip to content

The Checkpoint Sidecar Patch

This one is worth writing down because it's a bug you won't find in upstream, and it cost me a few build cycles to pin down.

The symptom

After a slot restore, the first request comes back at 0% cached. The second request is warm at 99%+. Everything looks healthy, the KV rehydrates perfectly, and then the first turn pays a full re-prefill anyway.

Root cause

slot.prompt.checkpoints live only in process memory. llama_state_seq_save_file never serializes them. So after SLOT_RESTORE the checkpoint list is empty, and the first request's rollback — the BPE re-tokenize of the prompt tail — finds no covering checkpoint:

forcing full prompt re-processing due to lack of cache data

The KV is there. The map telling the server which KV pages are valid for the prompt tail is not.

The fix

PR #206 on TheTom/llama-cpp-turboquant (apollo-mg, commit eaf98e612) persists checkpoints to a <state>.ckpt sidecar file (magic PKCL) at SLOT_SAVE and reloads them at SLOT_RESTORE. About 117 lines in one file.

It is not in upstream ggml-org master — I verified that on 2026-08-24. Bumping your pin will not pick it up. Tracking upstream: ggml-org PR #24028.

Results

A/B on a CPU rig, single slot, lfm2-5-1-2b: first-after-restore went from 0.00% to 99.48%. The patched log says it plainly:

saved 2 context checkpoints to sidecar
restored 2 context checkpoints to sidecar

On the prod box (llama-server-sidecar:2026.08.24-v4) the live test landed at 72–74% cached first-after-restore, which is where I'd expect it to sit with real agent traffic.

Gotchas applying it

Two things that bit me:

  1. The raw fork patch doesn't apply cleanly to a pin. Regenerate it with git apply --3way then git diff --cached.
  2. There's a semantic gap: the fork's parent has a token_count local that doesn't exist in my pin. Replace it with !slot->prompt.tokens.empty() or it won't compile.

Also, a separate issue still open: stale q4_0 .bin profile files left over from a cache-type flip fail restore with mismatched key type (8 != 2). Re-save all profile files after any cache-type change.

The image is locally built

The patched image is a locally built image, not the stock ghcr.io/ggml-org/llama.cpp:server-cuda. That matters because a compose file pointing at the stock image silently reverts the fix — first-after-restore cache goes back to 0% and nothing looks wrong. I hit exactly that when a branch pointed at the stock image while prod ran the patched one. Pin the sidecar tag, and check that the compose file matches what prod actually runs.

See also

Comments