Skip to content

The bug that wasn't in upstream

After a slot restore, the first request came back at 0% cached. The second one was warm at 99%. Nothing looked broken — the KV rehydrated fine, the logs were clean, and yet every restore paid a full re-prefill on the turn right after it.

I kept assuming it was a cache-type mismatch, because I'd had that problem before. It wasn't.

What the save function actually saves

llama_state_seq_save_file writes the KV pages. It does not write slot.prompt.checkpoints — those live only in process memory. So after a restore the checkpoint list is empty, and the first request's rollback (the BPE re-tokenize of the prompt tail) has nothing to match against:

forcing full prompt re-processing due to lack of cache data

The KV is there. What's missing is the map that tells the server which KV pages are valid for the prompt tail. Everything downstream of that is just the server being honest about not knowing.

The patch

PR #206 on TheTom/llama-cpp-turboquant (apollo-mg, commit eaf98e612) persists checkpoints to a <state>.ckpt sidecar file — magic bytes PKCL — at SLOT_SAVE, and reloads them at SLOT_RESTORE. About 117 lines in one file.

Warning

This is not in upstream ggml-org master. I checked on 2026-08-24. Bumping your pin will not pick it up. Upstream tracking is ggml-org PR #24028, still open.

What it bought

A/B on a CPU rig, single slot, lfm2-5-1-2b, same prompt, same everything:

first request after restore
unpatched 0.00% cached
patched 99.48% cached

The patched log says it plainly:

saved 2 context checkpoints to sidecar
restored 2 context checkpoints to sidecar

On the prod box the live test landed at 72–74% first-after-restore, which is where I'd expect it with real agent traffic rather than a clean A/B.

Two things that cost me build cycles

The raw fork patch does not apply to a pin. I had to regenerate it with git apply --3way and then git diff --cached. And there's a semantic gap: the fork's parent has a token_count local that doesn't exist in my pin, so it won't compile until you replace it with !slot->prompt.tokens.empty().

Three more bugs ate build cycles during the rollout — a null deref on slot.task in the fallback path, the SLT_WRN variadic macro, and patch-base drift. Also worth knowing: at the pinned commit the root CMakeLists sets CMAKE_RUNTIME_OUTPUT_DIRECTORY to ${CMAKE_BINARY_DIR}/bin, so the binary lands in /build/bin/llama-server, not /build/llama-server.

The part that bites you later

The patched image is a locally built image — llama-server-sidecar:2026.08.24-v4 — not the stock ghcr.io/ggml-org/llama.cpp:server-cuda. A compose file pointing at the stock image silently reverts the fix. First-after-restore cache goes back to 0% and nothing looks wrong, because the server still works.

I hit exactly that. A branch pointed at the stock image while prod ran the patched one, and a deploy would have quietly rolled the fix out. Now I diff the compose against what the box actually runs before deploying, and check the image tag matches.

One separate issue is still open on my side: stale q4_0 .bin profile files left over from a cache-type flip fail restore with mismatched key type (8 != 2). Re-save all profile files after any cache-type change.

Comments