Skip to content

Running Strata on a 5090: the setup I actually use

I've been running Qwen3.8-Flash-Next through Strata as the brain of my homelab for a while now, and the stack has settled into something worth writing down. 5090, 128 GB DDR5, IQ3_S quant, 260K context, 150-175 tok/s decode. It's the best local model I've used yet, and this is the configuration that got me there.

Start at IQ2_XS, move to IQ3_S when you feel the ceiling

I started with IQ2_XS. It felt smarter than a 27B dense model — which is the whole Flash-Next pitch, a 125B MoE with 6B active — but I hit thought loops on more complex problems. The model would get stuck cycling on itself during harder agentic work. Shifting to IQ3_S fixed it; I haven't seen a loop since. If you're on 32 GB of VRAM and 128 GB of RAM, IQ3_S fits comfortably and the quality difference is worth the swap.

The full reference numbers, the RAM gate, the Proxmox balloon trap, and the context-size math are on the Using Strata page — this post is about how it gets used, not how to install it.

Agents: Hermes and Opencode, all day and night

This box runs several Hermes agents and Opencode against it, multiple project channels, tool calling heavy. The engine handles agentic use excellently — I can run these boys all day and night without babysitting.

Where it surprised me is mobile development. In web development, Flash-Next and a 27B dense model are pretty close. But Flash is clearly better at mobile dev and especially emulator usage. The 27B struggled to actually drive Android/iOS apps in real-world usage; Flash cruises through mobile apps, testing out all the features. If your agent work touches device emulators, the MoE's world-model breadth shows up there in a way decode benchmarks won't tell you.

Hermes itself deserves the credit for the self-learning part. Ask it to generate a reading booklet from my first grader's monthly vocabulary list and it teaches itself the whole workflow the first time — vision-checking episode stills, matching text to frames, laying out the PDF — and just cranks it out the next time. It's a great homelab assistant and watchdog: daily reports on which external IP/ASN connections made it through the Pangolin proxy, Jellyfin usage, docker image versioning, backing up and updating docker stacks, watching PBS backups, querying Paperless. Pretty much anything you can reach over an API.

Concurrency: the fix landed, and it changed how I run things

I used to avoid concurrency on this engine entirely — two agents at long context would evict each other's prefix cache and both would crawl. That was a real architecture problem (one VRAM prefix window, no host-RAM tier), not a tuning problem. Last week's update fixed it: a conversation cache that parks whole snapshots in host RAM. My auxiliary code runners — the subagents Opencode spawns — now run excellent in parallel. In practice: a Hermes agent doing work, my wife chatting with our personal Hermes assistant, and somebody asking the Home Assistant voice assistant a question all resolve fine, without invalidating each other's cache like it used to. The measured before/after is on the Strata page (re-entry prefill drops about 4.4x).

And the piece I'm not running yet but want to: --shared-expert-arena, which shipped in v0.1.30. If you have multiple GPUs, you can run independent Strata lanes on each card — each with its own hot expert cache — all sharing the same copy of the experts in system RAM. A community reporter measured three lanes on three 5070 Tis at ~79 tok/s each, near-linear scaling, because the 50 GB expert arena only exists once in RAM. I've yet to test it, but it makes me want to throw another GPU in the box. (The multi-GPU section on the Strata page has the crossover data — lanes win for parallel agents, layer-split wins for one fast request, and PCIe x4 turns out to be fine.)

Dynamic VRAM resize: the agent stays alive through its own media job

This is the part of my setup I'd struggle to give up now. I run my own proxy in front of Strata — llama-slot-proxy — that handles GPU sharing between the LLM and ComfyUI, both in Docker on the same card. My agents queue their own media jobs: H3 videos, Flux and Qwen-Image renders, YuE2 music.

The old behavior was a full brain transplant: proxy stops Strata, starts ComfyUI, holds the GPU for the whole render, brings Strata back. The agent lost its engine for the duration and waited.

Strata changed that because its VRAM hold is elastic. The proxy can now shrink the hot expert cache in place — measured on my box, about 21 GB freed, cache down to ~1.5 GB while the render runs — hand the VRAM to ComfyUI, and regrow the cache when the job finishes. The agent stays resident and serving the whole time. It's noticeably slower mid-render, but it's alive: it watches its own job, polls its own queue, and when the media lands the proxy resizes it back to the full card with no reload. Small checkpoints (the klein family, SDXL, flux1-dev-fp8, z-image) route elastic; big ones (H3 video, fp8 Qwen-Image, Hunyuan 2.1) still stop/start the engine, because they genuinely can't co-reside on 32 GB. The size gate is a checkpoint-size check on the graph the agent submits, so the routing decision is automatic.

That elastic resize is a Strata property, not a proxy trick — it works because the expert cache is the thing holding VRAM, and it's the one structure that can be thrown away and refilled from the RAM tier without touching the model's weights or the conversation state. llama.cpp and NInfer don't have an equivalent knob; with those, the swap stays a swap.

The stack, in one paragraph

RTX 5090, 128 GB DDR5, Qwen3.8-Flash-Next IQ3_S at 260K context, Strata in Docker, llama-slot-proxy as the single keyed front door, ComfyUI and YuE2 sharing the card behind the swap gate, Hermes agents and Opencode as the clients, Home Assistant voice as the family's entry point. 150-175 tok/s, no cloud, one GPU doing double duty. Reference configs and the gotchas are on the Using Strata page.

Comments