DiggTop ⌕

Qwen3.8-27B on 21 GPUs: 42 Verified Runs

Qwen3.8-27B is the model our board has measured most — 42 verified rows across 21 different GPUs, 25 quantisation labels and 9 engines. That is a wide enough sample to answer a question a single review cannot: when the same model runs 26× faster on one card than another, how much of that is the card?

Most of it is not. The spread on a single GPU is as large as the spread between GPUs.

As of 2026-09-28. Every number below is an imported row that an upstream source verified — we did not run these tests. Each row carries engine, version, quant and context; the sources are linked at the end. Rows come from different people on different days, so read the columns before the ranking.

The board in one table

Best verified decode throughput per GPU, Qwen3.8-27B, with the format that produced it.

GPU Best tok/s Format / engine Context Rows
RTX 5090 435.1 IQ4_XS · lucebox 131K 6
Tesla V100 x4 391.0 NVFP4 · vLLM 1.2.2 32K 1
RTX PRO 6000 Blackwell 296.0 NVFP4 · SGLang 262K 3
CMP 170HX 255.9 W4A16 · vLLM 0.27.1 4.3K 4
RTX 3090 x2 219.8 AutoRound INT4 + FP8 KV · SGLang 0.5.19 262K 2
RTX 5090 x2 195.6 NVFP4 · SGLang dev 262K 1
RTX 3090 164.6 EXL3 3.5bpw · exllamav3 1.4.2 24K 6
RX 7900 XTX 154.2 MQ4 · hipfire 2K 2
Instinct MI100 x4 132.4 INT8 · vLLM 64K 1
RTX 3090 Ti 97.0 IQ4_XS · llama.cpp custom 204K 1
RTX 4090 141.2 int4 output-head-only · vLLM 38K 2
RTX 4090 D 70.9 Q4_K_M · llama.cpp 2.4K 2
Radeon AI Pro R9700 64.3 Q4_0 · llama.cpp b10711 262K 2
Intel Arc Pro B70 x2 101.9 AutoRound INT4 W4A16 · vLLM 0.20.2rc1 2K 2
RTX 5070 48.3 UD-Q2_K_XL · llama.cpp 32K 1
RTX 3080 44.4 IQ2_XXS · llama.cpp b10472 65K 1
RX 7800 XT 42.0 Unsloth Dynamic IQ4_XS · llama.cpp 121K 1
DGX Spark 56.0 NVFP4 W4A4 · SGLang dev-qwen38 262K 2
Tesla V100 x2 24.0 Q8_K_XL · llama.cpp 9463 131K 1
Intel Arc Pro B70 27.8 Q4_K_M · llama.cpp 8K 1
Radeon AI Pro R9700 x2 43.6 Q8_K_XL · llama.cpp 262K 1

Range: 16.8 → 435.1 tok/s. Nine engines. Twenty-five quant labels.

The same card, 26× apart

The RTX 5090's six rows are the clearest case:

Format Engine Context tok/s
IQ4_XS lucebox 131K 435.1
NVFP4 TensorRT-LLM (NInfer) 16K 213.0
W4A16 vLLM cu129-nightly 2K 137.9
W4A16 vLLM v0.27.1 2K 108.0
UD-Q6_K llama.cpp b10478 251K 79.4
FP8 vLLM 0.27.1 8K 16.8

A 26× spread on one card with one model. FP8 at 8K is the slowest row on the cheapest-to-reason-about configuration, and the same card's IQ4_XS row is the fastest in the entire 42-row set. Nothing about that is a property of the 5090.

The RX 7900 XTX shows the same effect at smaller scale: 154.2 tok/s with hipfire/MQ4 versus 58.0 with llama.cpp/Q4_K_M — 2.7× — on the same 24GB card, in the same month.

Practical read: if you are choosing a quantisation format and an engine for a 27B model, that choice moves throughput by more than upgrading the card does. Start there, then size the card for the context you need.

What a $1,300 used RTX 3090 actually delivers

The 3090 has the most rows (6) and the widest internal range of any card here:

Format Engine Context tok/s
EXL3 3.5bpw exllamav3 1.4.2 (Mia fork) 24K 164.6
EXL3 4.00bpw exllamav2 v1.5.0 8K 152.4
EXL3 4.00bpw buun-llama-cpp 262K 149.2
Q4_K_XL llama.cpp 0.3.0-dev 24K 81.9
EXL3 4.00bpw buun-llama-cpp 150K 27.8
EXL3 4.00bpw exllamav2 v1.5.0 150K 24.5

Used 3090s were $1,300 in our 2026-09-28 snapshot (up from $1,050 on 2026-09-04); a new RTX 5090 had reached $9,000 (up from $4,899.99 on 2026-09-10, as US retail stock dried up). The 3090's best row is 164.6 tok/s versus 213.0 for the 5090's best non-exotic row (NVFP4 via TensorRT-LLM) — a 1.3× throughput gap against a 6.9× price gap. The 5090's advantage is not raw decode here; it is the 32GB of VRAM, the NVFP4 path, and the 435 tok/s outlier that a purpose-built engine reached.

The context tax is the real cliff

The steepest drop in this dataset is not a card difference at all. One submitter's 3090 runs of the same EXL3 4.00bpw build:

Context tok/s
8K 152.4
150K 24.5

18× the context window for 6.2× less throughput — same GPU, same model, same quant, same engine version. Our earlier look at context scaling found that long context does not kill throughput on its own; what costs you is how the KV cache is quantised and how the engine schedules it. A 27B model at 128K+ is a memory-bandwidth problem, and the format you picked for the weights decides how much headroom is left.

Two caveats we are not going to paper over. First, a second submitter reported 149.2 tok/s at 262K on the same card and quant with a different llama.cpp fork — we cannot reconcile that with the 24.5 above, and both rows stay on the board with their source links. Second, these are 150K-row differences measured on rigs we did not control. Treat the shape as the finding, not the absolute numbers.

Old datacenter silicon is underrated

Three of the top ten rows are cards that cost a fraction of a modern consumer GPU, provided you can feed them:

The pattern is consistent: these parts have unusual memory configurations (HBM, 64GB+ per card) and the engine is doing the work. It also cuts the other way — a V100 x2 on Q8_K_XL managed 24.0 tok/s, slower than a single 3090, because the quant and engine mix did not suit the hardware.

What we would actually build for this model

If you have Run Expected
One 24GB card (3090, $1,300 used) EXL3 3.5–4bpw via exllamav3/exllamav2 150–165 tok/s at 8–24K
One 32GB card (5090, ~$9,000 new) NVFP4 via TensorRT-LLM or SGLang 195–213 tok/s, 16K+
Two 24GB cards AutoRound INT4 + FP8 KV via SGLang ~220 tok/s at 262K
A workstation budget RTX PRO 6000 Blackwell, 96GB 296 tok/s at 262K

A 48GB RTX A6000 was $5,975.06 new in the 2026-09-10 snapshot — more than two 3090s, for less throughput than two 3090s delivered on this model. VRAM is worth paying for; VRAM density at that price is not.

The board's Qwen3.8-27B profile has all 42 rows with their per-row context and source, and the GPU comparison puts two cards side by side per model. If you have a run we do not, submit it — the numbers here came from people who did.

FAQ

How much VRAM does Qwen3.8-27B need?

A 27–28B model fits in 24GB at 4-bit quantisation with room for a modest context window; our 3090 rows at 24K context are the proof. Above ~100K context, KV cache becomes the constraint and the same card drops from ~152 to ~25 tok/s.

Which GPU is fastest for Qwen3.8-27B?

The fastest verified row is an RTX 5090 at 435.1 tok/s with IQ4_XS on a custom engine, followed by a Tesla V100 x4 at 391.0 with NVFP4. Both results come from a specific format-plus-engine pairing, not from the card alone — the same 5090 runs the model at 16.8 tok/s with FP8 under vLLM.

Is quantisation or the GPU more important for throughput?

On this dataset, quantisation and engine choice move throughput more than a tier of GPU does: 16.8 → 435.1 tok/s on one RTX 5090, versus 164.6 → 213.0 tok/s between the best 3090 and the best non-exotic 5090 row.

How do you verify these numbers?

Every row is imported from a source link and carries hardware, quant, engine version and context — the four items our methodology treats as the minimum for a benchmark to mean anything. Rows missing any of them stay out of the board.

Sources

↑