Skip to content

Using llama.cpp

llama.cpp is still the workhorse here. It's the engine I trust most because it's the one that gives you control over the cache, and the cache is what decides whether an agent feels fast or feels broken.

The box is an RTX 5090 (32GB) running as a QEMU VM, CUDA 13.2, CPU host-passthrough. llama.cpp runs as server-cuda in router mode, which is what lets me swap models without touching Docker.

Router mode

Router mode means the server holds a preset file and serves whatever model you load into it:

# models-preset.ini
[qwen3.6-27b]
model = /models/qwen3.6-27b.gguf
alias = qwen3.6-27b
ctx = 131072
cache_type_k = q4_0
cache_type_v = q4_0
mtp = /models/mtp-draft.gguf

Start the server with --models-preset models-preset.ini, then swap models through the API instead of restarting the container:

curl -X POST http://192.168.4.29:8080/models/load -d '{"model_name":"qwen3.6-27b"}'
curl -X POST http://192.168.4.29:8080/models/unload -d '{"model_name":"qwen3.6-27b"}'

A model swap costs VRAM, not a container rebuild, so a media job or a second engine can take the GPU and give it back without anyone restarting anything.

Slots, save and restore

The router keeps slots. Saving and restoring a slot is what makes a long conversation survive a GPU swap:

curl -X POST http://192.168.4.29:8080/slots/1 \
  -d '{"action":"save","model":"qwen3.6-27b"}'

Note the model field — it's not optional. In router mode the server routes by model name, so a save or restore without it has no idea which model the slot belongs to. This bit me once because the swap logic used to hardcode the --model-name flag, and after a ComfyUI round-trip it would reload the default model and bind the slot to the wrong thing.

Cache

--cache-ram gives llama.cpp a host-RAM checkpoint tier. Without it, the prefix cache is one window in VRAM and two agents working at once evict each other every turn. With it, a miss costs a read from RAM instead of a full re-prefill.

This is why llama.cpp still beats the fancier engines on multi-agent traffic, even when its decode numbers are worse. I covered the measurement in Strata — the same workload, 5.4% miss rate on llama.cpp against 10.4% on an engine with no host-RAM tier.

The checkpoint sidecar patch

There's one bug worth knowing about, and it's not in upstream. See the sidecar patch.

See also

Comments