DiggTop ⌕

One 180B MoE, Five Formats: 34 Verified Runs

For most of September, Qwen3.8-Flash-Next was not one model on our board. It was seven entries: a base page, four format-suffixed pages (-GGUF, -NVFP4, -GSQ-RCO-IQ3XXS, -W4A16-FP8PLE), a -REAP-256-duo-GGUF page and a -hibrid48 page, each holding the rows for a single quantisation lineage. The board has merged them — the four format pages now 301-redirect into the merged profile, and -REAP-256-duo-GGUF redirects into the surviving REAP page — which makes this the first piece that can put the formats side by side on the same 180B mixture-of-experts model.

The merged page carries 34 verified rows, 11 hardware configurations, 4 engines and 15 quant labels, spanning 20.3 to 474 tok/s — a 23× range on one model.

As of 2026-09-28. Every row below is an imported, upstream-verified measurement; we did not run these tests. Each row carries hardware, quant, engine version and context, and the sources are linked at the end. Rows come from different people on different days on different rigs, so read the columns before the ranking.

What merged, and what did not

Former entity Where it went Verified rows
Qwen3.8-Flash-Next (base) merged profile 31
Qwen3.8-Flash-Next-GGUF redirects into base —
Qwen3.8-Flash-Next-NVFP4 redirects into base —
Qwen3.8-Flash-Next-GSQ-RCO-IQ3XXS redirects into base —
Qwen3.8-Flash-Next-W4A16-FP8PLE redirects into base —
Qwen3.8-Flash-Next-REAP-256-duo-GGUF REAP-256-duo page 1
Qwen3.8-Flash-Next-hibrid48 hibrid48 page 2

The redirects are not cosmetic: format identity now lives in the quant column, where it belongs. But two of the three surviving entities are genuinely different checkpoints, not quantisations of one file — 31 rows are the base model, 1 is a 256-expert-pruned REAP build, 2 are the hibrid48 mixed-precision build. We compare them below; we do not treat them as interchangeable. This is the same trap we flagged when 3.5 bits beat 4 bits.

Five format families, one model

Grouping the board's 15 quant labels into the five format families the merged page now carries:

Format family Quant labels on the board Rows Best tok/s Worst tok/s Engines
NVFP4 family NVFP4, NVFP4+FP8 hybrid, NVFP4 (output head) 15 474 37.4 vllm, sglang
W4A16 + FP8 W4A16-FP8PLE, W4A16+FP8-KV 2 186.9 145.8 vllm
EXL3 EXL3, EXL3 2.50bpw, EXL3 3.05bpw 8 72.0 20.9 exllamav2
GGUF I-quants IQ3XXS, IQ3_S, IQ4_XS, UD-IQ1_S, UD-IQ3_XXS, Unsloth-Dynamic-IQ3_XXS 8 96.1 20.3 llama.cpp
REAP-256-duo GGUF Unsloth-Dynamic-Q3_K_XL 1 53.6 53.6 llama.cpp

EXL3 is the odd one out: it never had a split page, so its rows were always filed under the base name. That is part of why the merge was worth doing — the format the board was quietly measuring most consistently was the one with the least visible identity.

The two fastest rows, and the pair we publish both sides of

The single highest number is 474 tok/s: the hibrid48 checkpoint on 2× DGX Spark, vLLM 0.29, 262,144-token context. The same checkpoint on a single DGX Spark is recorded at 60 tok/s with vLLM 0.29.0 at 262,000 — 7.9× apart. Different card count, different build, same suffix. We cannot attribute that gap to the format, so both rows stay on the board with their sources.

The fastest row that is not a variant checkpoint is 186.9 tok/s on 2× CMP 170HX with W4A16-FP8PLE under vLLM 0.29.0 at 262,052 context, on two repurposed mining cards at $1,050 each in the current snapshot.

On the base checkpoint, NVFP4 is not one speed either — it is a spread:

Base-checkpoint NVFP4 row Engine Context tok/s
DGX Spark ×2 vLLM + patches 262,144 80.0
DGX Spark ×1 SGLang 8874c51a 262,144 55.0
DGX Spark ×1 vLLM 8a728663 262,144 43.9
DGX Spark ×1 vLLM hybrid build, 327 ctx 327 37.4

Same format, same-ish hardware, 2.1× between the fastest and slowest row. Engine and build are doing as much work here as the format is.

What actually fits on consumer hardware

Config Format Engine Context tok/s Peak VRAM
RTX 3090 ×1 EXL3 2.50bpw exllamav2 262,144 38.6 —
RTX 3090 ×1 EXL3 2.50bpw exllamav2 175,000 20.9 —
CMP 170HX ×1 UD-IQ1_S llama.cpp pr-27742 65,536 46.1 43.18 GB
RTX 3090 ×2 Unsloth-Dynamic-IQ3_XXS llama.cpp 217,088 40.1 42.85 GB
RTX 3090 ×2 Unsloth-Dynamic-Q3_K_XL (REAP) llama.cpp 32,768 53.6 33.90 GB
Radeon AI Pro R9700 ×1 IQ4_XS llama.cpp 32,768 25.6 32.91 GB
Intel Arc Pro B70 ×2 IQ3_S llama.cpp 16,384 20.3 —

A second conflict we are not smoothing over: the two RTX 3090 EXL3 2.50bpw rows come from one submitter, one engine build, one quant — and the 262,144-token run (38.6) is faster than the 175,000-token run (20.9). Longer context should not decode faster. Both rows stay.

The structural finding is in the VRAM column. A 180B MoE only runs on 24–48GB of consumer memory at roughly 1.5–3 bits per weight: UD-IQ1_S peaked at 43.18 GB on a single 64GB card, and Unsloth-Dynamic-IQ3_XXS peaked at 42.85 GB across two 24GB cards while still holding a 217K context. Pruning is the other lever — the REAP-256-duo build, which drops 256 experts, peaked at 33.90 GB on those same two 3090s, about 9GB less, and ran 13 tok/s faster at a much shorter 32K window.

Note what is missing: the RTX 5090 — the only 32GB card in our 2026-09-28 price snapshot — has no verified rows for this model at all, and neither does any other RTX 50-series part. The rows above come from RTX 3090s, a repurposed CMP 170HX, and RDNA/Arc workstation cards.

The price of a format, at the 2026-09-28 snapshot

Config Best tok/s Cards Cost tok/s per dollar (derived)
CMP 170HX ×2 186.9 2 × $1,050 used $2,100 0.089
RTX 3090 ×4 145.8 4 × $1,300 used $5,200 0.028
RTX 3090 ×1 38.6 1 × $1,300 used $1,300 0.030
RTX 3090 ×2 (REAP) 53.6 2 × $1,300 used $2,600 0.021
Intel Arc Pro B70 ×2 20.3 2 × $949.99 new $1,899.98 0.011

Every price above is a 2026-09-28 snapshot figure; the last column is our arithmetic on the two preceding columns, not a measurement. The CMP 170HX pair returns 3.2× the tok/s per dollar of the four-3090 build — with the caveat that the 170HX is a repurposed mining card with no display outputs, awkward drivers, and HBM that the price source's own headline frames as a memory-unlock story.

Two price conflicts worth publishing rather than averaging. Used RTX 3090s were $1,050 in our 2026-09-04 snapshot and are $1,300 on 2026-09-28 — the earlier figure is the one quoted in our 27B article. The RTX 5090 is a wider gap: $4,899.99 for a Newegg listing on 2026-09-10 versus $9,000 as a US street average on 2026-09-28. Different markets, so the $9,000 is a market signal rather than a checkout price — but it is the number the current snapshot carries, and the 5090 has no benchmark rows here to justify either figure.

What we would actually run

If you have Run Expected
One used 24GB card ($1,300) EXL3 2.50bpw via exllamav2 20–39 tok/s, context up to 262K
One 64GB ex-mining card ($1,050) UD-IQ1_S GGUF via llama.cpp ~46 tok/s at 64K, 43GB peak
2× used 24GB, short context REAP-256-duo Q3_K_XL ~54 tok/s at 32K, 34GB peak
2× used 24GB, long context Unsloth-Dynamic-IQ3_XXS ~40 tok/s at 217K, 43GB peak
2× CMP 170HX (~$2,100) W4A16-FP8PLE via vLLM 0.29+ ~187 tok/s at 256K — best value here
4× RTX 3090 (~$5,200) W4A16+FP8-KV via vLLM nightly ~146 tok/s at 256K

The benchmark board has all 34 rows with per-row context and source; how we verify explains the four fields a row needs to be published at all; submit a run if you have one we do not. The price reports carry the snapshot these figures came from. For the two card pairings above that matter most, see CMP 170HX vs RTX 3090 and DGX Spark vs RTX 3090.

FAQ

Which quantisation format is fastest for Qwen3.8-Flash-Next?

NVFP4-family kernels, but only on specific hardware: 474 tok/s on 2× DGX Spark with the hibrid48 build, and 186.9 tok/s on 2× CMP 170HX with W4A16-FP8PLE under vLLM. On consumer 24GB cards the question inverts — you are choosing whichever sub-4-bit format fits at all, and GGUF I-quants or EXL3 are the realistic answers.

Can a 180B MoE run on a single 24GB GPU?

Yes, at low bit width. One RTX 3090 ran Qwen3.8-Flash-Next with EXL3 2.50bpw at 38.6 tok/s with a 262,144-token context, and 20.9 tok/s at 175,000 from the same submitter and engine build. The longer run being the faster one is a contradiction we publish rather than explain away.

How much VRAM does Qwen3.8-Flash-Next need?

Weights land in the 33–44GB band in our rows: UD-IQ1_S peaked at 43.18GB at 64K context, Unsloth-Dynamic-IQ3_XXS at 42.85GB across two 3090s at 217K, and the pruned REAP-256-duo build at 33.90GB at 32K. KV cache sits on top of those figures, which is why the long-context rows need the most memory.

Why does the same model show 474 tok/s and 60 tok/s on DGX Spark?

Because the row identity is not just the card. The 474 figure is the hibrid48 checkpoint on 2× DGX Spark with vLLM 0.29; the 60 figure is the same checkpoint on a single DGX Spark with vLLM 0.29.0. Card count and build both changed, so we keep both rows visible instead of picking one.

How do you verify these numbers?

Every row is imported from a source link and carries hardware, quant, engine version and context — the minimum our methodology requires. These are upstream-verified measurements; DiggTop did not run them.

Sources

Imported upstream measurements as of 2026-09-28. DiggTop did not run these tests; each row carries its engine version, quantisation and context, and passed import review under the methodology rules.

↑