Hardware comparison: Qwen3.8-Flash-Next and 27B¶
Two models, six machines, one question: where does each model actually run fast. Every number on this page is a published community measurement with a link, or a number I measured on my own boxes. Where nobody has published a number, the cell says so instead of guessing.
The short version: for Qwen3.8-Flash-Next, decode speed is set by memory bandwidth and by which tensor your engine decides to page to SSD. That second part is not a footnote — it's a 14x swing, and it's the difference between a Mac that runs this model at 160 tok/s and the same class of Mac running it at 9.
Why this model behaves differently than everything else¶
Qwen3.8-Flash-Next is a 125B MoE with only 6B active per token, plus a 51B n-gram embedding table and a 4B MTP head (model card, architecture paper). Three facts drive every number below:
- Decode reads the active parameters, not the total. 6B active at ~4-bit is roughly 2.6-3.4 GB of weight traffic per token, so decode ceilings are high on any fast-memory box. A dense 70B at 4-bit reads ~40 GB per token and crawls on the same hardware.
- The n-gram table is a cache, and caches are happy to be slow. A token reads two rows of it (~5 KB). It can live on NVMe via mmap and cost a few percent. This is what Strata does on my 5090, what MTPLX does on every Mac since 2.12.0 (the table ships as a sidecar file), and what the Strix Halo recipe does with 47.7 GiB of table streamed from SSD.
- The experts must NOT be the thing you page. Each token routes through 10 experts x 48 layers, a different ten every token. Page those to SSD and decode collapses.
The 9 tok/s trap
A community oMLX benchmark on a 96 GB M5 Ultra shows 9.0 tok/s for this model. The settings JSON explains it: moe_expert_offload_enabled: true, resident_fraction: 0.25 — the 112 GB table-resident MLX build didn't fit, so the engine kept the n-gram table hot and paged three-quarters of the experts to SSD instead. GPU utilization averaged 19%. Same silicon, right split (Bare Speed pack, ~74 GB resident, table on SSD), and the expected speed is an order of magnitude higher. When a Flash-Next number looks absurd, check which tensor is being paged before believing the hardware verdict. The same trap, expert-pruning variant, is documented in this 128 GB Mac writeup.
Memory requirements by quant (Flash-Next)¶
| Build | Download | Resident in RAM | Notes |
|---|---|---|---|
| GGUF UD-IQ1_S | 72.5 GB | ~45-50 GB + table mmap | table quantized harder at 1-bit tiers (Unsloth) |
| GGUF UD-IQ3_XXS | 82 GB | ~60 GB + table mmap | fits 96 GB Macs / 64 GB RAM + GPU hybrid |
| GGUF UD-Q3_K_XL | 90 GB | ~65 GB + table mmap | |
| GGUF UD-IQ4_XS | 93.7 GB | ~68 GB + table mmap | tight on 96 GB |
| GGUF UD-Q4_K_XL | 111.4 GB | ~80 GB + table mmap | 128 GB tier |
| Strata IQ2_XS / IQ3_XXS / IQ3_S | ~68 GB | RAM gate: 44 / 56 / 58 GiB visible | expert tiering: hot experts VRAM, all experts host RAM (Strata) |
| MLX oQ4 (table resident) | 112 GB | 112 GB | 128 GB tier; this is the build that fails on 96 GB |
| MTPLX Bare Speed | 106.3 GB | ~74 GB, table on SSD | (pack card) |
| MTPLX Optimized Speed | 115.1 GB | ~83 GB, table on SSD | recommended |
| MTPLX Optimized Quality (8-bit) | 170 GB | ~128.5 GiB, table on SSD | 256 GB tier |
Qwen3.8-Flash-Next: decode and prefill by machine¶
| Machine | Memory / bandwidth | Engine | Decode (tok/s) | Prefill |
|---|---|---|---|---|
| DGX Spark (GB10) | 128 GB / 273 GB/s | vLLM NVFP4 + MTP | 41.7 median, 42-45 code, 27.4 prose; 17 without MTP | 1,666 tok/s at 32K |
| DGX Spark | 128 GB | llama.cpp Q4_K_XL | 20.5-23.4 | — |
| Ryzen AI Halo (Strix Halo, 8060S) | 128 GB / ~256 GB/s | EngramHalo.cpp + MTP | 82 code (3-bit-class quant), 55 JSON, ~22 prose | slow: 200K prompt took 17.3 min to first token |
| MacBook Pro M5 Max | 128 GB / ~600 GB/s | MTPLX 2.12 + MTP | 54.6-122.4 on real OpenCode tasks, 125.8 best request; 68.4 at 16K, 60.9 at 100K | 1,453 tok/s at 4K, 1,094 at 65K |
| Mac Studio M5 Pro | 128 GB | — | no published numbers | — |
| Mac Studio M5 Ultra (64c) | 96 GB / 1.2 TB/s | oMLX (experts paged — wrong split) | 9.0 (see trap above) | 90.9 tok/s (also degraded) |
| Mac Studio M5 Ultra (64c) | 96 GB | MTPLX Bare Speed (right split) | not yet benchmarked; ~74 GB resident fits | — |
| Mac Studio M5 Ultra (80c) | 256 GB / 1.2 TB/s | oMLX recipe + lookup drafts | 158-165 fresh, 225-272 agent edit turns | 4,759 tok/s at 32K |
| Mac Studio M5 Ultra (80c) | 256 GB | TensorFold 0.3.6.1 | 183-197 fresh, 437 edit turns (quality cost: 5.5x KLD) | 2,826 tok/s at 32K |
| Mac Studio M5 Ultra (80c) | 256 GB | oMLX stock 0.7.0rc1 | 81.2 at 8K code (MTP on); MacStories: 52 at 16K, 111.6 prose | 4,755 tok/s at 8K; 2,887 at 16K |
| RTX 5090 (my box) | 32 GB VRAM + host RAM | Strata IQ2_XS | 129-172 (measured, strata page); 150-200 on current build with IQ3_S | 64K prompt in 11.1 s (~5,800 tok/s) |
| RTX 3090 | 24 GB VRAM + 64 GB RAM | Strata | ~100-140 (README estimate, ±20%; K8V4 KV measured 99 tok/s at 198K) | ~490-520 tok/s (older engine) |
| RTX 5070 | 12 GB VRAM + 64 GB DDR5 | Strata | 93.5 Q2_0, 53 IQ3_S (measured, DETAILS.md) | 2,653 tok/s at 32K (Q2_0, engine 0.1.36) |
Sources: ai-muninn Spark benchmark, NVIDIA forum Spark thread, two-Spark results, Strix Halo recipe + results, MTPLX measurements, M5 Ultra oMLX recipe + engine comparison, oMLX community benchmarks (80c run, 96c run), MacStories M5 Ultra review via the LLMCheck audit, mlx-serve engine comparison issue.
Qwen3.8-27B (dense): decode and prefill by machine¶
| Machine | Memory | Engine | Decode (tok/s) | Prefill |
|---|---|---|---|---|
| DGX Spark | 128 GB | vLLM NVFP4 + MTP | ~28 single-stream (recipe + harness); up to 77.3 with tuned speculative decoding across an 11-config study | — |
| Ryzen AI Halo | 128 GB | — | no direct 27B number published; dense 70B at 4-bit runs ~5 tok/s on this class of bandwidth, which is the honest anchor | — |
| MacBook Pro M5 Max | 128 GB | MLX | ~29-30 (LLMCheck estimate, labeled estimate) | — |
| Mac Studio M5 Ultra | 96 GB | — | fits at 4-bit (~16 GB) with room to spare; no direct 96 GB measurement, bandwidth is the same as the 256 GB model | — |
| Mac Studio M5 Ultra | 256 GB | LM Studio (plain decode) | ~55 (BGR review); LLMCheck's bandwidth estimate of 57 landed within 4% | 1,701 tok/s (MacStories) |
| Mac Studio M5 Ultra | 256 GB | oMLX + MTP, long prompts | 48 (long 8-16K prompts slow decode even with MTP) | — |
| RTX 5090 (my box) | 32 GB | NInfer nvfp4 | 184-204 (measured, NInfer page) | slower than Strata at long context |
| RTX 3090 | 24 GB | llama.cpp / NInfer | no recorded benchmark | — |
Anecdotal configs worth knowing¶
Beyond the measured tables, the Strata community has pushed this model onto a lot of ordinary hardware. These are single reports, not benchmarks — treat them as anecdotal — but they show where the expert-tiering design goes:
- RTX 3060 12 GB + 48 GB RAM (Ryzen 5 3600): Q2_0 ran at ~44 tok/s — a 125B-class model on a 12 GB card from 2021 (writeup).
- 64 GB RAM + 16 GB VRAM: 40-50 tok/s at 256K context reported (summary).
- 64 GB RAM at 256K context with IQ3_S: users confirmed it runs with RAM to spare (Strata issue #406).
- RTX 5090 with the NVFP4 fork: 141 -> 158 tok/s just from widening the MTP draft vocabulary (Strata DETAILS.md).
- RTX 5080 + IQ3_S and an RTX 2070 (Turing) both contributed measured runs to the repo — the engine supports RTX 20/30/40/50 and a long list of Radeon cards (AMD_HIP.md).
- Dual-GPU: Strata doesn't split one model across two cards, but a second card can take the vision encoder (
cuda_devicein the vision section, DETAILS.md). I've seen 2x3090 and 2x5090 combinations discussed in community threads for other engines; nobody has published a Flash-Next number for them that I could find.
The pattern across every one of these: the binding constraint is RAM capacity for the expert pool, and the second is VRAM for the hot-expert cache. A fast CPU with lots of RAM keeps a small GPU productive.
What the tables actually say¶
Bandwidth sets the decode ceiling, but engines move 2-3x on the same chip. On the M5 Max, oMLX measured 33 tok/s and MTPLX measured 68-125 on the same model in the same month. On the Spark, the same checkpoint went 17 -> 41.7 tok/s with the MTP recipe and no weight changes. The published tok/s is always an engine-plus-config number; the chip is only half of it.
MTP is the single biggest lever, and it's workload-shaped. Prose drafts well (predictable), code drafts worse, and agent edit turns — where the output copies the input — are the jackpot: sethforprivacy's prompt-lookup drafts hit 272 tok/s on edit turns because the continuation is already sitting in the context. My Strata setup runs the same trick with MTP on the 5090.
The M5 Ultra's headline is prefill, not decode. MacStories measured decode +40% over M3 Ultra but prefill +150-250% (16K prompt: 1,143 -> 2,887 tok/s, TTFT 13.9s -> 5.6s), from the Neural Accelerators. For agents that re-read long contexts, that's the gain that shows up in wall-clock time.
96 GB is the awkward tier for Flash-Next. The table-resident MLX builds don't fit and page experts, which is how you get 9 tok/s. The GGUF and MTPLX splits (table on SSD, experts hot) do fit, and the arithmetic says they run well, but nobody has published that benchmark yet. 128 GB is the floor where every engine's fast path works.
My 5090 with Strata is still the fastest decode box on this page. 150-200 tok/s on the current build, and the NInfer measurements show what 1.8 TB/s of HBM buys. The Mac's case was never raw decode — it's 96-512 GB of resident room, no slot swapping, and prefill speed.
Related¶
- Using Strata — the expert-tiering engine behind my 5090 numbers
- Using NInfer · Using llama.cpp · llama-slot-proxy