Using NInfer¶
NInfer is the engine that runs my day-to-day agent traffic. It's a CUDA-only inference engine, and on this box it serves the 27B model as an nvfp4 artifact.
What's actually running¶
The production container is ninfer-eval, compose-managed from ~/docker/ninfer/docker-compose.yml, port 8095. The launch args:
ninfer-serve --model /models/qwen3_8_27b_nvfp4.ninfer \
--max-context 260000 --kv-capacity 260000 --kv-dtype fp8 \
--spec mtp --draft-tokens 3 --lm-head-draft --preserve-thinking
MTP spec-decoding is what gives NInfer its decode numbers — 184–204 t/s on this box, which is faster than llama.cpp and faster than the 125B MoE. NInfer wins short prompts outright.
The artifact format trap¶
This is the thing I'd want anyone to know before touching NInfer. Upstream broke the artifact format between v2 and v3. My source tree (~/ninfer-src) is deliberately pinned to match the v2 artifact, and you should never fast-forward it — a rebuild from that tree produces an engine that cannot read our model.
The v3 line lives in a separate worktree (~/ninfer-src-v3) with a converted artifact. The v3 artifact was verified with upstream's own reader — version 3, 1190 objects, exactly the count the upgrade script lists as correct for qwen3.8-27b/nvfp4. But the v3 engine has never loaded it on a GPU, so it stays a candidate, not the deployed fallback.
The swap trap¶
strata-swap.sh to-ninfer only runs docker start ninfer-eval. It does not apply a compose file. So the container comes up with whatever image and model path it was last created with, and editing a .yml on disk changes nothing until you recreate.
Before trusting any swap, ask what it will actually run:
A recreate is GPU-free — create makes the container and leaves it stopped — so staging a candidate costs nothing until the swap itself touches VRAM.
Both pairs sit on disk, both artifacts are 23.7GB, so switching is a container recreate, not a copy:
| pair | image | artifact | status |
|---|---|---|---|
| v2 | ninfer:local |
qwen3_8_27b_nvfp4.ninfer |
deployed fallback — has actually run on the GPU |
| v3 | ninfer:v3 |
qwen3_8_27b_nvfp4_v3.ninfer |
built, never started |
The rule I follow: the deployed pair is the verified pair. A build that has never run on hardware is not a fallback, no matter how correct its metadata looks.
Why NInfer can't coexist with Strata¶
Each engine wants about 31GB of VRAM on a 32GB card. There's no room for both, so the box runs one engine at a time and the swap script handles the handover. The llama-slot-proxy is what makes that workable from the client side.