Skip to content

Using Strata

Strata is a purpose-built engine for Qwen3.8-Flash-Next. The idea is expert tiering: hot experts on the GPU, all experts in host RAM, with MTP draft spec-decoding and chunked prefill. That's how a 125B MoE fits on a 32GB card.

Repo is Niko1221/Strata (MIT), CUDA-only. On this box it serves qwen3.8-flash-next-iq2_xs on port 8082 with a Bearer key.

Design that's worth copying

The image is disposable and everything real lives on host volumes:

  • /opt/strata = host ~/docker/strata/app — the git clone, the venv, the compiled engine
  • /data = host ~/docker/strata/data — models, packs, MTP

No model and no engine in the image. So a container rebuild costs nothing, and the 68GB model survives. That's the right shape for something you'll be rebuilding often.

Two things that will bite you on first run:

  1. --host 0.0.0.0 is mandatory. setup.py start() doesn't pass --host to server.py; the bind comes from the config, which defaults to 127.0.0.1, and the container port mapping breaks.
  2. Pre-compile the engine before first serve. It's CPU-only and saves 10–20 minutes at first request. But compile_driver.py calls setup.gpus() → nvidia-smi, so it must run with --gpus all or it asserts "no usable GPU found".

The RAM gate

setup.py refuses to start when visible RAM is below the quant's requirement (measured from /proc/meminfo MemTotal):

quant needs
Q2_0 / IQ2_XS ≥44 GiB visible
IQ3_XXS ≥56 GiB
IQ3_S ≥58 GiB

The pitfall here is the Proxmox balloon. I watched this VM drop from 48.0 GiB to 15.9 GiB MemTotal on the same boot — virtio_balloon loaded, "smallest model needs 32 GB", setup stops. The fix is host-side: pin the VM's memory so the host can't reclaim it. My swap script now preflights RAM and aborts before touching anything else. Don't remove that guard.

Measured numbers (5090, IQ2_XS, 128K int8, 256 tokens, temp 0)

ctx Strata cold TTFT / decode NInfer cold TTFT / decode
1k 536ms / 129 t/s 336ms / 184 t/s
16k 3039ms / 135 t/s 5542ms / 204 t/s
64k 11065ms / 172 t/s 31849ms / 193.75 t/s

Read that properly: NInfer wins short prompts. Strata's 64k prefill is 2.9x faster, and its decode climbs with context until it ties. So Strata is a long-context engine, not a fast one.

Cold start is not what the docs claim. First start after a model download took ~690s to reach health; a warm restart ~280s. Plan for 5–12 minutes.

Context size costs less than you think

KV lives in VRAM, so context and expert capacity trade off directly. At int8, 12 cells per token (12 full-attention layers) × 1056 B/cell = 12,672 B/token:

context KV in VRAM
128K 1.55 GiB
260K 3.09 GiB

--expert-cache auto sizes from whatever VRAM is left. Measured, 128K vs 260K: experts in VRAM 17,822 → 16,427, cache 23.93 → 22.02 GiB, hit rate 99.3–99.7% → 97.8–99.1%, throughput within noise (two of three measurements were faster at 260K). The capacity you give up is the least-used tail of the profile-filled cache. Don't assume a linear penalty.

The multi-agent problem, and what fixed it

The worst thing about Strata for me was total cache invalidation when agents work at the same time. There's one VRAM prefix window and no host-RAM tier, so two agents at 164k and 190k tokens have a working set larger than the window — mutual eviction is guaranteed, not a tuning problem. Miss rate 10.4% against llama.cpp's 5.4%, and llama.cpp wins only because it has --cache-ram.

Upstream fixed this in v0.1.30 with a conversation cache that parks whole snapshots in host RAM:

--conversation-cache-mib 8192
--conversation-cache-slots 4
--conversation-cache-min-free-mib 2560

Measured with two alternating 57k/60k conversations:

turn prompt cached prefill wall
A1 cold 57,122 0 9,170 ms 9.3 s
B1 cold 59,721 0 10,034 ms 10.1 s
A2 (the old guaranteed re-read) 57,124 49,152 2,098 ms 2.2 s
B2 59,723 49,152 2,259 ms 2.4 s

evictions=0 with both conversations live. Prefill on re-entry drops about 4.4x, restore costs 75ms. Sizing measured: ~1.5 GB per 57k-token conversation — about 2x the int8 KV, because a snapshot carries running state, checkpoints, GDN/SSM state and draft KV.

Two caveats I'd state honestly. Reuse resumes from a checkpoint (taken every 16K tokens and at each assistant turn), so the tail after the boundary is re-read — that's the 2 seconds. And issue #342 is still open upstream: successive turns of the same conversation each consume a slot, so a long subagent chain can push the parent out. Alternating different chats works today; delegation chains don't.

RAM is the binding constraint on this box. With 48GB total and ~8GB available while serving, the 8 GiB example budget gets refused by the 2560 MiB floor. Expect 4–5 GiB, roughly two 57k chats. And the balloon risk applies again — a live MemTotal shrink makes parking get silently skipped. Safe degradation, but the cache stops helping.

Engine flags via env, not config

setup.py rebuilds cfg["args"] from scratch on every container start, so anything hand-edited into strata-<model>.json is silently lost on the next docker compose up -d. The only durable way in is an env hook that appends to cfg["args"] before the config is written. That's how the parking flags get in.

Related: you cannot install over a running engine. build_engine() ends with shutil.copy2 onto the executing binary → Text file busy. Compile first (no downtime), then stop the container, then install.

Serving quirks

  • Reasoning model: thinking arrives in message.reasoning_content, the answer in message.content. A low max_tokens gives you content: null with finish_reason: "length".
  • /v1/models requires the API key.
  • Serves one request at a time. No concurrency.

See also

Comments