One 180B MoE, Five Formats: 34 Verified Runs
For most of September, Qwen3.8-Flash-Next was not one model on our board. It was seven entries: a base page, four format-suffixed pages (-GGUF, -NVFP4, -GSQ-RCO-IQ3XXS, -W4A16-FP8PLE), a -REAP-256-duo-GGUF page and a -hibrid48 page, each holding the rows for a single quantisation lineage. The board has merged them — the four format pages now 301-redirect into the merged profile, and -REAP-256-duo-GGUF redirects into the surviving REAP page — which makes this the first piece that can put the formats side by side on the same 180B mixture-of-experts model.
The merged page carries 34 verified rows, 11 hardware configurations, 4 engines and 15 quant labels, spanning 20.3 to 474 tok/s — a 23× range on one model.
As of 2026-09-28. Every row below is an imported, upstream-verified measurement; we did not run these tests. Each row carries hardware, quant, engine version and context, and the sources are linked at the end. Rows come from different people on different days on different rigs, so read the columns before the ranking.
What merged, and what did not
| Former entity | Where it went | Verified rows |
|---|---|---|
Qwen3.8-Flash-Next (base) |
merged profile | 31 |
Qwen3.8-Flash-Next-GGUF |
redirects into base | — |
Qwen3.8-Flash-Next-NVFP4 |
redirects into base | — |
Qwen3.8-Flash-Next-GSQ-RCO-IQ3XXS |
redirects into base | — |
Qwen3.8-Flash-Next-W4A16-FP8PLE |
redirects into base | — |
Qwen3.8-Flash-Next-REAP-256-duo-GGUF |
REAP-256-duo page | 1 |
Qwen3.8-Flash-Next-hibrid48 |
hibrid48 page | 2 |
The redirects are not cosmetic: format identity now lives in the quant column, where it belongs. But two of the three surviving entities are genuinely different checkpoints, not quantisations of one file — 31 rows are the base model, 1 is a 256-expert-pruned REAP build, 2 are the hibrid48 mixed-precision build. We compare them below; we do not treat them as interchangeable. This is the same trap we flagged when 3.5 bits beat 4 bits.
Five format families, one model
Grouping the board's 15 quant labels into the five format families the merged page now carries:
| Format family | Quant labels on the board | Rows | Best tok/s | Worst tok/s | Engines |
|---|---|---|---|---|---|
| NVFP4 family | NVFP4, NVFP4+FP8 hybrid, NVFP4 (output head) | 15 | 474 | 37.4 | vllm, sglang |
| W4A16 + FP8 | W4A16-FP8PLE, W4A16+FP8-KV | 2 | 186.9 | 145.8 | vllm |
| EXL3 | EXL3, EXL3 2.50bpw, EXL3 3.05bpw | 8 | 72.0 | 20.9 | exllamav2 |
| GGUF I-quants | IQ3XXS, IQ3_S, IQ4_XS, UD-IQ1_S, UD-IQ3_XXS, Unsloth-Dynamic-IQ3_XXS | 8 | 96.1 | 20.3 | llama.cpp |
| REAP-256-duo GGUF | Unsloth-Dynamic-Q3_K_XL | 1 | 53.6 | 53.6 | llama.cpp |
EXL3 is the odd one out: it never had a split page, so its rows were always filed under the base name. That is part of why the merge was worth doing — the format the board was quietly measuring most consistently was the one with the least visible identity.
The two fastest rows, and the pair we publish both sides of
The single highest number is 474 tok/s: the hibrid48 checkpoint on 2× DGX Spark, vLLM 0.29, 262,144-token context. The same checkpoint on a single DGX Spark is recorded at 60 tok/s with vLLM 0.29.0 at 262,000 — 7.9× apart. Different card count, different build, same suffix. We cannot attribute that gap to the format, so both rows stay on the board with their sources.
The fastest row that is not a variant checkpoint is 186.9 tok/s on 2× CMP 170HX with W4A16-FP8PLE under vLLM 0.29.0 at 262,052 context, on two repurposed mining cards at $1,050 each in the current snapshot.
On the base checkpoint, NVFP4 is not one speed either — it is a spread:
| Base-checkpoint NVFP4 row | Engine | Context | tok/s |
|---|---|---|---|
| DGX Spark ×2 | vLLM + patches | 262,144 | 80.0 |
| DGX Spark ×1 | SGLang 8874c51a |
262,144 | 55.0 |
| DGX Spark ×1 | vLLM 8a728663 |
262,144 | 43.9 |
| DGX Spark ×1 | vLLM hybrid build, 327 ctx | 327 | 37.4 |
Same format, same-ish hardware, 2.1× between the fastest and slowest row. Engine and build are doing as much work here as the format is.
What actually fits on consumer hardware
| Config | Format | Engine | Context | tok/s | Peak VRAM |
|---|---|---|---|---|---|
| RTX 3090 ×1 | EXL3 2.50bpw | exllamav2 | 262,144 | 38.6 | — |
| RTX 3090 ×1 | EXL3 2.50bpw | exllamav2 | 175,000 | 20.9 | — |
| CMP 170HX ×1 | UD-IQ1_S | llama.cpp pr-27742 |
65,536 | 46.1 | 43.18 GB |
| RTX 3090 ×2 | Unsloth-Dynamic-IQ3_XXS | llama.cpp | 217,088 | 40.1 | 42.85 GB |
| RTX 3090 ×2 | Unsloth-Dynamic-Q3_K_XL (REAP) | llama.cpp | 32,768 | 53.6 | 33.90 GB |
| Radeon AI Pro R9700 ×1 | IQ4_XS | llama.cpp | 32,768 | 25.6 | 32.91 GB |
| Intel Arc Pro B70 ×2 | IQ3_S | llama.cpp | 16,384 | 20.3 | — |
A second conflict we are not smoothing over: the two RTX 3090 EXL3 2.50bpw rows come from one submitter, one engine build, one quant — and the 262,144-token run (38.6) is faster than the 175,000-token run (20.9). Longer context should not decode faster. Both rows stay.
The structural finding is in the VRAM column. A 180B MoE only runs on 24–48GB of consumer memory at roughly 1.5–3 bits per weight: UD-IQ1_S peaked at 43.18 GB on a single 64GB card, and Unsloth-Dynamic-IQ3_XXS peaked at 42.85 GB across two 24GB cards while still holding a 217K context. Pruning is the other lever — the REAP-256-duo build, which drops 256 experts, peaked at 33.90 GB on those same two 3090s, about 9GB less, and ran 13 tok/s faster at a much shorter 32K window.
Note what is missing: the RTX 5090 — the only 32GB card in our 2026-09-28 price snapshot — has no verified rows for this model at all, and neither does any other RTX 50-series part. The rows above come from RTX 3090s, a repurposed CMP 170HX, and RDNA/Arc workstation cards.
The price of a format, at the 2026-09-28 snapshot
| Config | Best tok/s | Cards | Cost | tok/s per dollar (derived) |
|---|---|---|---|---|
| CMP 170HX ×2 | 186.9 | 2 × $1,050 used | $2,100 | 0.089 |
| RTX 3090 ×4 | 145.8 | 4 × $1,300 used | $5,200 | 0.028 |
| RTX 3090 ×1 | 38.6 | 1 × $1,300 used | $1,300 | 0.030 |
| RTX 3090 ×2 (REAP) | 53.6 | 2 × $1,300 used | $2,600 | 0.021 |
| Intel Arc Pro B70 ×2 | 20.3 | 2 × $949.99 new | $1,899.98 | 0.011 |
Every price above is a 2026-09-28 snapshot figure; the last column is our arithmetic on the two preceding columns, not a measurement. The CMP 170HX pair returns 3.2× the tok/s per dollar of the four-3090 build — with the caveat that the 170HX is a repurposed mining card with no display outputs, awkward drivers, and HBM that the price source's own headline frames as a memory-unlock story.
Two price conflicts worth publishing rather than averaging. Used RTX 3090s were $1,050 in our 2026-09-04 snapshot and are $1,300 on 2026-09-28 — the earlier figure is the one quoted in our 27B article. The RTX 5090 is a wider gap: $4,899.99 for a Newegg listing on 2026-09-10 versus $9,000 as a US street average on 2026-09-28. Different markets, so the $9,000 is a market signal rather than a checkout price — but it is the number the current snapshot carries, and the 5090 has no benchmark rows here to justify either figure.
What we would actually run
| If you have | Run | Expected |
|---|---|---|
| One used 24GB card ($1,300) | EXL3 2.50bpw via exllamav2 | 20–39 tok/s, context up to 262K |
| One 64GB ex-mining card ($1,050) | UD-IQ1_S GGUF via llama.cpp | ~46 tok/s at 64K, 43GB peak |
| 2× used 24GB, short context | REAP-256-duo Q3_K_XL | ~54 tok/s at 32K, 34GB peak |
| 2× used 24GB, long context | Unsloth-Dynamic-IQ3_XXS | ~40 tok/s at 217K, 43GB peak |
| 2× CMP 170HX (~$2,100) | W4A16-FP8PLE via vLLM 0.29+ | ~187 tok/s at 256K — best value here |
| 4× RTX 3090 (~$5,200) | W4A16+FP8-KV via vLLM nightly | ~146 tok/s at 256K |
The benchmark board has all 34 rows with per-row context and source; how we verify explains the four fields a row needs to be published at all; submit a run if you have one we do not. The price reports carry the snapshot these figures came from. For the two card pairings above that matter most, see CMP 170HX vs RTX 3090 and DGX Spark vs RTX 3090.
FAQ
Which quantisation format is fastest for Qwen3.8-Flash-Next?
NVFP4-family kernels, but only on specific hardware: 474 tok/s on 2× DGX Spark with the hibrid48 build, and 186.9 tok/s on 2× CMP 170HX with W4A16-FP8PLE under vLLM. On consumer 24GB cards the question inverts — you are choosing whichever sub-4-bit format fits at all, and GGUF I-quants or EXL3 are the realistic answers.
Can a 180B MoE run on a single 24GB GPU?
Yes, at low bit width. One RTX 3090 ran Qwen3.8-Flash-Next with EXL3 2.50bpw at 38.6 tok/s with a 262,144-token context, and 20.9 tok/s at 175,000 from the same submitter and engine build. The longer run being the faster one is a contradiction we publish rather than explain away.
How much VRAM does Qwen3.8-Flash-Next need?
Weights land in the 33–44GB band in our rows: UD-IQ1_S peaked at 43.18GB at 64K context, Unsloth-Dynamic-IQ3_XXS at 42.85GB across two 3090s at 217K, and the pruned REAP-256-duo build at 33.90GB at 32K. KV cache sits on top of those figures, which is why the long-context rows need the most memory.
Why does the same model show 474 tok/s and 60 tok/s on DGX Spark?
Because the row identity is not just the card. The 474 figure is the hibrid48 checkpoint on 2× DGX Spark with vLLM 0.29; the 60 figure is the same checkpoint on a single DGX Spark with vLLM 0.29.0. Card count and build both changed, so we keep both rows visible instead of picking one.
How do you verify these numbers?
Every row is imported from a source link and carries hardware, quant, engine version and context — the minimum our methodology requires. These are upstream-verified measurements; DiggTop did not run them.
Sources
- Merged model entity and all 34 rows: Qwen3.8-Flash-Next profile, benchmark board, pulled 2026-09-28.
hibrid48474 @262,144 on DGX Spark ×2, vLLM 0.29 — x.com/vr8vr8/status/2099167079492432229. Same checkpoint 60 @262,000 on DGX Spark ×1, vLLM 0.29.0 — x.com/vr8vr8/status/2099416741478568221.- CMP 170HX ×2 W4A16-FP8PLE 186.9 @262,052, vLLM 0.29.0 — x.com/MoonlitMaven/status/2102471011882975451.
- RTX 3090 ×4 W4A16+FP8-KV 145.8 @262,144, vLLM nightly (qwen4exp-ple) — x.com/Tech2Wild/status/2096754559263617169.
- Tesla V100 ×2 IQ3XXS 96.1 @4K, 87.5 @512, 80.9 @131K, llama.cpp
b11030-mix-5ff778e— x.com/PeasantSmith/status/2102804368286175514. - DGX Spark NVFP4 base checkpoint 80 @262,144 on ×2 (vLLM + patches) — x.com/vr8vr8/status/2096338499876082054; 55 @262,144 on ×1 (SGLang
8874c51a) — x.com/shantanugoel/status/2100191973290480127; 43.9 @262,144 (vLLM8a728663) — x.com/Tech2Wild/status/2101524927639601296; NVFP4+FP8 hybrid 37.4–44.2 across nine contexts — x.com/redp314/status/2096641228821405850. - DGX Spark EXL3 72 @240K, 69.8 @4K, 69.5 @128K — x.com/ViC305/status/2100741339042427140; EXL3 3.05bpw 67.6 @75K, 59 @32K, 56 @1K — x.com/WescheNex1q/status/2100984588617031740.
- RTX 3090 EXL3 2.50bpw 38.6 @262,144 and 20.9 @175,000 — x.com/mr_r0b0t/status/2100570088025801099.
- RTX 3090 ×2 REAP-256-duo Q3_K_XL 53.6 @32,768, 33.90GB peak — localmaxxing run cmtdbfo0g003mlm01o3htzqao; Unsloth-Dynamic-IQ3_XXS 40.1 @217,088, 42.85GB peak — localmaxxing run cmtdbd9xn0031lm01jszmlexx.
- CMP 170HX UD-IQ1_S 46.1 @65,536, 43.18GB peak — localmaxxing run cmtauavsx000yqq01durnd750.
- Radeon AI Pro R9700 IQ4_XS 25.6 @32,768, 32.91GB peak — localmaxxing run cmtbqjvpj003bqq01595qg0db; R9700 ×3 UD-IQ3_XXS 28.1 @829 — localmaxxing run cmtddazgh003rlm015xk99gyj; Intel Arc Pro B70 ×2 IQ3_S 20.3 @16,384 — localmaxxing run cmtigo39n04e2p401eumx5joo.
- Prices, all snapshot 2026-09-28: used CMP 170HX $1,050 (street report) — overclockers.ua; used RTX 3090 $1,300 (eBay resale) — gamebastion.com; new Intel Arc Pro B70 $949.99 (Newegg) — techpowerup.com; new RTX 5090 $9,000 (US street average, ixbt 2026-09-15) — ixbt.com. Prior snapshots: used RTX 3090 $1,050 on 2026-09-04 — howardresourcegroup.com; new RTX 5090 $4,899.99 on 2026-09-10 (Newegg listing) — Newegg Gigabyte GV-N5090GAMING-OC-32GD. See the price reports.
Imported upstream measurements as of 2026-09-28. DiggTop did not run these tests; each row carries its engine version, quantisation and context, and passed import review under the methodology rules.