Why I send 429 instead of 503
When a media job owns the GPU, the proxy has to tell the client to go away for a few minutes. The status code it uses to do that matters more than I expected.
When a media job owns the GPU, the proxy has to tell the client to go away for a few minutes. The status code it uses to do that matters more than I expected.
This one applies well beyond the box it happened on, so I'm writing it down separately.
I had a swap script that started a container with docker start. I'd edited the compose file to point at a new image and a new model path, ran the swap, and it came up running the old pair. Nothing was wrong with the compose file. Nothing was wrong with the script. The container was simply doing what containers do.
The slot proxy fronts the GPU box, and my own LLM calls route through it. So the moment I deploy the proxy, I'm deploying the thing that's currently serving me.
The llama slot proxy is the piece I'd least expect someone to build and the piece I'd least like to work without. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to.