Using Strata¶
Strata is a purpose-built engine for Qwen3.8-Flash-Next. The idea is expert tiering: hot experts on the GPU, all experts in host RAM, with MTP draft spec-decoding and chunked prefill. That's how a 125B MoE fits on a 32GB card.
Repo is Niko1221/Strata (MIT), CUDA-only. On this box it serves qwen3.8-flash-next-iq2_xs on port 8082 with a Bearer key.
Design that's worth copying¶
The image is disposable and everything real lives on host volumes:
/opt/strata= host~/docker/strata/app— the git clone, the venv, the compiled engine/data= host~/docker/strata/data— models, packs, MTP
No model and no engine in the image. So a container rebuild costs nothing, and the 68GB model survives. That's the right shape for something you'll be rebuilding often.
Two things that will bite you on first run:
--host 0.0.0.0is mandatory.setup.py start()doesn't pass--hosttoserver.py; the bind comes from the config, which defaults to127.0.0.1, and the container port mapping breaks.- Pre-compile the engine before first serve. It's CPU-only and saves 10–20 minutes at first request. But
compile_driver.pycallssetup.gpus()→nvidia-smi, so it must run with--gpus allor it asserts "no usable GPU found".
The RAM gate¶
setup.py refuses to start when visible RAM is below the quant's requirement (measured from /proc/meminfo MemTotal):
| quant | needs |
|---|---|
| Q2_0 / IQ2_XS | ≥44 GiB visible |
| IQ3_XXS | ≥56 GiB |
| IQ3_S | ≥58 GiB |
The pitfall here is the Proxmox balloon. I watched this VM drop from 48.0 GiB to 15.9 GiB MemTotal on the same boot — virtio_balloon loaded, "smallest model needs 32 GB", setup stops. The fix is host-side: pin the VM's memory so the host can't reclaim it. My swap script now preflights RAM and aborts before touching anything else. Don't remove that guard.
Measured numbers (5090, IQ2_XS, 128K int8, 256 tokens, temp 0)¶
| ctx | Strata cold TTFT / decode | NInfer cold TTFT / decode |
|---|---|---|
| 1k | 536ms / 129 t/s | 336ms / 184 t/s |
| 16k | 3039ms / 135 t/s | 5542ms / 204 t/s |
| 64k | 11065ms / 172 t/s | 31849ms / 193.75 t/s |
Read that properly: NInfer wins short prompts. Strata's 64k prefill is 2.9x faster, and its decode climbs with context until it ties. So Strata is a long-context engine, not a fast one.
Cold start is not what the docs claim. First start after a model download took ~690s to reach health; a warm restart ~280s. Plan for 5–12 minutes.
Context size costs less than you think¶
KV lives in VRAM, so context and expert capacity trade off directly. At int8, 12 cells per token (12 full-attention layers) × 1056 B/cell = 12,672 B/token:
| context | KV in VRAM |
|---|---|
| 128K | 1.55 GiB |
| 260K | 3.09 GiB |
--expert-cache auto sizes from whatever VRAM is left. Measured, 128K vs 260K: experts in VRAM 17,822 → 16,427, cache 23.93 → 22.02 GiB, hit rate 99.3–99.7% → 97.8–99.1%, throughput within noise (two of three measurements were faster at 260K). The capacity you give up is the least-used tail of the profile-filled cache. Don't assume a linear penalty.
The multi-agent problem, and what fixed it¶
The worst thing about Strata for me was total cache invalidation when agents work at the same time. There's one VRAM prefix window and no host-RAM tier, so two agents at 164k and 190k tokens have a working set larger than the window — mutual eviction is guaranteed, not a tuning problem. Miss rate 10.4% against llama.cpp's 5.4%, and llama.cpp wins only because it has --cache-ram.
Upstream fixed this in v0.1.30 with a conversation cache that parks whole snapshots in host RAM:
Measured with two alternating 57k/60k conversations:
| turn | prompt | cached | prefill | wall |
|---|---|---|---|---|
| A1 cold | 57,122 | 0 | 9,170 ms | 9.3 s |
| B1 cold | 59,721 | 0 | 10,034 ms | 10.1 s |
| A2 (the old guaranteed re-read) | 57,124 | 49,152 | 2,098 ms | 2.2 s |
| B2 | 59,723 | 49,152 | 2,259 ms | 2.4 s |
evictions=0 with both conversations live. Prefill on re-entry drops about 4.4x, restore costs 75ms. Sizing measured: ~1.5 GB per 57k-token conversation — about 2x the int8 KV, because a snapshot carries running state, checkpoints, GDN/SSM state and draft KV.
Two caveats I'd state honestly. Reuse resumes from a checkpoint (taken every 16K tokens and at each assistant turn), so the tail after the boundary is re-read — that's the 2 seconds. And issue #342 is still open upstream: successive turns of the same conversation each consume a slot, so a long subagent chain can push the parent out. Alternating different chats works today; delegation chains don't.
RAM is the binding constraint on this box. With 48GB total and ~8GB available while serving, the 8 GiB example budget gets refused by the 2560 MiB floor. Expect 4–5 GiB, roughly two 57k chats. And the balloon risk applies again — a live MemTotal shrink makes parking get silently skipped. Safe degradation, but the cache stops helping.
Engine flags via env, not config¶
setup.py rebuilds cfg["args"] from scratch on every container start, so anything hand-edited into strata-<model>.json is silently lost on the next docker compose up -d. The only durable way in is an env hook that appends to cfg["args"] before the config is written. That's how the parking flags get in.
Related: you cannot install over a running engine. build_engine() ends with shutil.copy2 onto the executing binary → Text file busy. Compile first (no downtime), then stop the container, then install.
Serving quirks¶
- Reasoning model: thinking arrives in
message.reasoning_content, the answer inmessage.content. A lowmax_tokensgives youcontent: nullwithfinish_reason: "length". /v1/modelsrequires the API key.- Serves one request at a time. No concurrency.
See also¶
- llama-slot-proxy — how the box fronts all three engines
- NInfer · Using llama.cpp