Why I send 429 instead of 503
When a media job owns the GPU, the proxy has to tell the client to go away for a few minutes. The status code it uses to do that matters more than I expected.
When a media job owns the GPU, the proxy has to tell the client to go away for a few minutes. The status code it uses to do that matters more than I expected.
This one applies well beyond the box it happened on, so I'm writing it down separately.
I had a swap script that started a container with docker start. I'd edited the compose file to point at a new image and a new model path, ran the swap, and it came up running the old pair. Nothing was wrong with the compose file. Nothing was wrong with the script. The container was simply doing what containers do.
A 125B MoE reported better decode than the 27B it replaced, and it felt slower. Both things were true. The engine log wasn't lying — it was answering a different question.
The slot proxy fronts the GPU box, and my own LLM calls route through it. So the moment I deploy the proxy, I'm deploying the thing that's currently serving me.
After a slot restore, the first request came back at 0% cached. The second one was warm at 99%. Nothing looked broken — the KV rehydrated fine, the logs were clean, and yet every restore paid a full re-prefill on the turn right after it.
The llama slot proxy is the piece I'd least expect someone to build and the piece I'd least like to work without. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to.
A 125B MoE needs about 44 GiB of visible RAM to start. On the same boot, the VM went from 48.0 GiB to 15.9 GiB, and setup stopped with "smallest model needs 32 GB".
Three requests hit exactly 65,536 completion tokens. 240, 252 and 287 seconds each. Thirteen minutes — 12% of a 1.75 hour window — in three requests, none of which delivered an answer.
I ran this for a while with two agents on the same engine and both of them felt sluggish, and I couldn't work out why. The engine reported good numbers. The experience didn't match.
The answer is the prefix cache, and it's structural — not a tuning knob.
In this post I'm going to discuss project and user agents for each of the platforms as well as native and third party applications that allow remote access to your agents.