Skip to content

2026

A container runs what it was last created with

This one applies well beyond the box it happened on, so I'm writing it down separately.

I had a swap script that started a container with docker start. I'd edited the compose file to point at a new image and a new model path, ran the swap, and it came up running the old pair. Nothing was wrong with the compose file. Nothing was wrong with the script. The container was simply doing what containers do.

The bug that wasn't in upstream

After a slot restore, the first request came back at 0% cached. The second one was warm at 99%. Nothing looked broken — the KV rehydrated fine, the logs were clean, and yet every restore paid a full re-prefill on the turn right after it.

A front layer for local inference

The llama slot proxy is the piece I'd least expect someone to build and the piece I'd least like to work without. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to.

Runaway thinking

Three requests hit exactly 65,536 completion tokens. 240, 252 and 287 seconds each. Thirteen minutes — 12% of a 1.75 hour window — in three requests, none of which delivered an answer.

Why two agents on one local LLM feel slow

I ran this for a while with two agents on the same engine and both of them felt sluggish, and I couldn't work out why. The engine reported good numbers. The experience didn't match.

The answer is the prefix cache, and it's structural — not a tuning knob.