Skip to content

Why I send 429 instead of 503

When a media job owns the GPU, the proxy has to tell the client to go away for a few minutes. The status code it uses to do that matters more than I expected.

It answers:

429 { "code": "media_inflight" }
Retry-After: <real estimate>

The estimate is per-owner full generation time minus elapsed time since the swap, clamped to [30, 900] seconds — about 240s for music, 180s for ComfyUI.

Why not 503

Because of how the client classifies errors. Hermes maps 503 to overloaded, which ignores Retry-After and falls back to a 2/4/8 second jittered backoff. For a swap that owns the GPU for three to ten minutes, that backoff exhausts long before the engine comes back, and an agent with no configured fallback errors out mid-generation.

429 maps to rate_limit, which honors Retry-After (capped at 600s). One clean retry lands on the restore.

The actual difference

Same information, different status code, different behaviour. 503 says "I'm busy" and the client retries on a timer that's too short. 429 says "retry at this time" and the client waits. The header is only useful if the classifier reads it.

The swap itself

The GPU can't hold an inference engine and a PyTorch media container at the same time — each wants about 31 GB on a 32 GB card. So the gate stops the chat container, runs the media job, and brings the engine back.

Two things I got wrong first:

  • The media container must fully stop, not just /free. An idle PyTorch context holds about 500 MB, and the engine's startup reservation has ~100 MB of slack. Not enough.
  • The gate has to follow the selected backend's container, readiness path, and readiness budget. It used to hardcode the NInfer container, so while a different engine held the GPU, a media job stopped a container that wasn't running, never freed the VRAM, died on the allocation, and the "restore" resurrected the wrong engine on top of the one that was actually serving.

Readiness differs per engine. One engine's /v1/models 401s without a key, so the gate probes /health. And the budget: one engine reloads in 280–690 seconds measured. A flat 60–120 second budget aborts a healthy restore and strands the GPU on the media container. The other engine restores in about 20 seconds.

That's why I swap to the fast engine for batches of media work. A media job while the slow engine holds the card costs 5–12 minutes of LLM downtime.

And the container has to exist

The gate restarts the chat container with docker start, so it must exist in a stopped state. docker compose stop is correct; docker compose down deletes it and the restore has nothing to start.

Comments