A front layer for local inference¶
The llama slot proxy is the piece I'd least expect someone to build and the piece I'd least like to work without. It sits on port 8081 and fronts whatever engine currently owns the GPU, and it's the only thing the clients talk to.
Without it, every consumer would need to know which engine is up, which model is loaded in it, and whether a media job just took the card. That's three things that change during the day, and every client would need to handle all three.
What it does¶
Profiles by key. Each client gets a key, and the key identifies a profile — model, personality, context limits. My agent runs as roy, the others have their own. The key is local: it identifies the profile at the proxy, it isn't forwarded upstream by default.
A backend registry. Every upstream is declared once and one is selected by name:
SLOT_PROXY_BACKENDS={"ninfer":{"url":"http://192.168.4.29:8095","key":""},
"strata":{"url":"http://192.168.4.29:8082","key":"<api-key>","mode":"ninfer","container":"strata"},
"llamacpp":{"url":"http://192.168.4.29:8080","key":""}}
SLOT_PROXY_BACKEND=strata
The mode field is what bit me. Personality is derived from the selected name, so SLOT_PROXY_BACKEND=strata alone gives you the llama.cpp full slot scheduler — GPU swap, admission control — which is wrong for an engine that already owns the GPU and serves one resident model. mode: ninfer borrows the pass-through personality instead.
Metering. Every request lands in a SQLite DB with prompt_tokens, cached_tokens, prompt_ms, predicted_ms, duration_ms, ttft_ms, ttft_visible_ms, reasoning_chars, finish_reason, served_model.
Tip
This is the part that paid for itself. Every performance conclusion I've drawn about the box came out of querying this DB, not out of the engine log. The engine log reports decode tokens per second sampled mid-request; the DB reports what actually happened per request. Those two numbers disagreed, and the DB was right.
Admin surface. /policy for the effective merged policy, /slots/{id}/evict to forcibly release a holder, /slots/{id}/lock as an ephemeral eviction-protection flag, plus /gpu, /slots, /inactive.
Why a proxy and not just an engine¶
Three reasons, all of them came from hitting the problem:
- Engines can't coexist on one card. Each wants about 31 GB on a 32 GB card. So the box runs one at a time, and something has to know which one. That's the proxy.
- The GPU is shared with media work. ComfyUI and the music model need the card too. The proxy gates the swap and answers with a real
Retry-Afterwhile the media job runs. - Per-request data has to be recorded somewhere the client can't skip. A client that misbehaves — burning 65,536 tokens on thinking, or evicting another agent's cache — is only visible if something in the middle writes it down.
Dashboard¶
/data returns a dict (rows / total / has_more), not a list. Pagination walks newest to oldest with zero overlap, and the key filter scopes the total to that key and works combined with recent_before. The slot timeline is one row per slot with a profile color legend above it — bars colored per profile, not one row per profile.
Repo: Wildium/llama-slot-proxy. Dashboard at llama.thomasjwilde.com.
The write-ups on the specific traps: deploying a proxy to itself · why 429 and not 503 · the metering numbers.