Skip to content

Using NInfer

NInfer is the engine that runs my day-to-day agent traffic. It's a CUDA-only inference engine, and on this box it serves the 27B model as an nvfp4 artifact.

What's actually running

The production container is ninfer-eval, compose-managed from ~/docker/ninfer/docker-compose.yml, port 8095. The launch args:

ninfer-serve --model /models/qwen3_8_27b_nvfp4.ninfer \
  --max-context 260000 --kv-capacity 260000 --kv-dtype fp8 \
  --spec mtp --draft-tokens 3 --lm-head-draft --preserve-thinking

MTP spec-decoding is what gives NInfer its decode numbers — 184–204 t/s on this box, which is faster than llama.cpp and faster than the 125B MoE. NInfer wins short prompts outright.

The artifact format trap

This is the thing I'd want anyone to know before touching NInfer. Upstream broke the artifact format between v2 and v3. My source tree (~/ninfer-src) is deliberately pinned to match the v2 artifact, and you should never fast-forward it — a rebuild from that tree produces an engine that cannot read our model.

The v3 line lives in a separate worktree (~/ninfer-src-v3) with a converted artifact. The v3 artifact was verified with upstream's own reader — version 3, 1190 objects, exactly the count the upgrade script lists as correct for qwen3.8-27b/nvfp4. But the v3 engine has never loaded it on a GPU, so it stays a candidate, not the deployed fallback.

The swap trap

strata-swap.sh to-ninfer only runs docker start ninfer-eval. It does not apply a compose file. So the container comes up with whatever image and model path it was last created with, and editing a .yml on disk changes nothing until you recreate.

Before trusting any swap, ask what it will actually run:

docker inspect ninfer-eval --format '{{.Config.Image}} {{.Config.Cmd}}'

A recreate is GPU-free — create makes the container and leaves it stopped — so staging a candidate costs nothing until the swap itself touches VRAM.

Both pairs sit on disk, both artifacts are 23.7GB, so switching is a container recreate, not a copy:

pair image artifact status
v2 ninfer:local qwen3_8_27b_nvfp4.ninfer deployed fallback — has actually run on the GPU
v3 ninfer:v3 qwen3_8_27b_nvfp4_v3.ninfer built, never started

The rule I follow: the deployed pair is the verified pair. A build that has never run on hardware is not a fallback, no matter how correct its metadata looks.

Why NInfer can't coexist with Strata

Each engine wants about 31GB of VRAM on a 32GB card. There's no room for both, so the box runs one engine at a time and the swap script handles the handover. The llama-slot-proxy is what makes that workable from the client side.

See also

Comments