llama-slot-proxy¶
This is what makes the box usable day to day. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to. Without it, every consumer would need to know which engine is up, which model is loaded, and whether a media job just took the card.
Repo: Wildium/llama-slot-proxy. Dashboard at llama.thomasjwilde.com.
What it actually does¶
Profiles by key. Each client gets a key, and the key identifies a profile — model, personality, context limits. My agent runs as roy, the other agents have their own. The key is local: it identifies the profile at the proxy, it isn't forwarded upstream by default.
Backend registry. Every upstream is declared once, and one is selected by name:
SLOT_PROXY_BACKENDS={"ninfer":{"url":"http://192.168.4.29:8095","key":""},
"strata":{"url":"http://192.168.4.29:8082","key":"<api-key>","mode":"ninfer","container":"strata"},
"llamacpp":{"url":"http://192.168.4.29:8080","key":""}}
SLOT_PROXY_BACKEND=strata
The mode field matters more than it looks. Personality is derived from the selected name, so SLOT_PROXY_BACKEND=strata alone gives you the llama.cpp full slot scheduler — GPU swap, admission control — which is wrong for an engine that already owns the GPU and serves one resident model. "mode": "ninfer" borrows the pass-through personality instead.
Metering. Every request lands in a SQLite DB (usage-data/usage.db) with prompt_tokens, cached_tokens, prompt_ms, predicted_ms, duration_ms, ttft_ms, reasoning_chars, finish_reason, served_model. That DB is the source of truth for per-request latency — not the engine log. Every performance conclusion I've drawn about this box came out of querying it.
The media swap gate¶
The GPU can't hold an engine and a PyTorch media container at the same time, so the gate stops the chat container, runs the media job, and brings the chat engine back.
Two things people get wrong:
- The media container must fully stop, not just
/free. An idle PyTorch context holds about 500MB, and the engine's startup reservation has ~100MB of slack. That's not enough. - The gate has to follow the selected backend's container, readiness path, and readiness budget. It used to hardcode the NInfer container, so while Strata held the GPU a media job stopped
ninfer-eval— a no-op — never freed Strata's 31GB, died on VRAM, and the restore resurrected NInfer on top of Strata.
Readiness differs per engine. Strata's /v1/models 401s without a key, so the gate probes /health. And the budget: Strata reloads in 280–690 seconds measured, so a flat 60–120s budget aborts a healthy restore and strands the GPU on the media container. NInfer restores in about 20 seconds. A media job while Strata holds the card costs 5–12 minutes of LLM downtime, which is why I swap to NInfer for batches of media work.
The chat container must exist — stopped, not removed — because the gate restarts it with docker start. So the swap script uses docker compose stop, never down.
Why 429 and not 503¶
When a media job owns the GPU, the proxy answers 429 {code: media_inflight} with a Retry-After that's a real estimate — per-owner full generation minus elapsed since the swap, clamped to [30, 900].
That was a deliberate choice. Hermes' error classifier maps 503 to overloaded, which ignores Retry-After and falls back to a 2/4/8s jittered backoff. For a swap that owns the GPU for 3–10 minutes, that exhausts long before the engine restores, and a local agent with no fallback errored out mid-generation. 429 maps to rate_limit, which honors Retry-After (capped at 600s), so one clean retry lands on the restore.
Deploying it is a special case¶
The proxy serves the agent that deploys the proxy. If I deploy it the normal way, I cut my own head off mid-request. So deploys are delayed and self-safe: the deploy is scheduled to land after the current request completes, not during it.
Dashboard notes¶
/data returns a dict (rows/total/has_more), not a list. Pagination walks newest to oldest with zero overlap, and the key filter scopes the total to that key and works combined with recent_before. The slot timeline is one row per slot with a profile color legend above it — bars colored per profile, not one row per profile.