Qwen3.8-27B on 21 GPUs: 42 Verified Runs
Qwen3.8-27B is the model our board has measured most — 42 verified rows across 21 different GPUs, 25 quantisation labels and 9 engines. That is a wide enough sample to answer a question a single review cannot: when the same model runs 26× faster on one card than another, how much of that is the card?
Most of it is not. The spread on a single GPU is as large as the spread between GPUs.
As of 2026-09-28. Every number below is an imported row that an upstream source verified — we did not run these tests. Each row carries engine, version, quant and context; the sources are linked at the end. Rows come from different people on different days, so read the columns before the ranking.
The board in one table
Best verified decode throughput per GPU, Qwen3.8-27B, with the format that produced it.
| GPU | Best tok/s | Format / engine | Context | Rows |
|---|---|---|---|---|
| RTX 5090 | 435.1 | IQ4_XS · lucebox |
131K | 6 |
| Tesla V100 x4 | 391.0 | NVFP4 · vLLM 1.2.2 | 32K | 1 |
| RTX PRO 6000 Blackwell | 296.0 | NVFP4 · SGLang | 262K | 3 |
| CMP 170HX | 255.9 | W4A16 · vLLM 0.27.1 | 4.3K | 4 |
| RTX 3090 x2 | 219.8 | AutoRound INT4 + FP8 KV · SGLang 0.5.19 | 262K | 2 |
| RTX 5090 x2 | 195.6 | NVFP4 · SGLang dev | 262K | 1 |
| RTX 3090 | 164.6 | EXL3 3.5bpw · exllamav3 1.4.2 | 24K | 6 |
| RX 7900 XTX | 154.2 | MQ4 · hipfire |
2K | 2 |
| Instinct MI100 x4 | 132.4 | INT8 · vLLM | 64K | 1 |
| RTX 3090 Ti | 97.0 | IQ4_XS · llama.cpp custom | 204K | 1 |
| RTX 4090 | 141.2 | int4 output-head-only · vLLM | 38K | 2 |
| RTX 4090 D | 70.9 | Q4_K_M · llama.cpp | 2.4K | 2 |
| Radeon AI Pro R9700 | 64.3 | Q4_0 · llama.cpp b10711 | 262K | 2 |
| Intel Arc Pro B70 x2 | 101.9 | AutoRound INT4 W4A16 · vLLM 0.20.2rc1 | 2K | 2 |
| RTX 5070 | 48.3 | UD-Q2_K_XL · llama.cpp | 32K | 1 |
| RTX 3080 | 44.4 | IQ2_XXS · llama.cpp b10472 | 65K | 1 |
| RX 7800 XT | 42.0 | Unsloth Dynamic IQ4_XS · llama.cpp | 121K | 1 |
| DGX Spark | 56.0 | NVFP4 W4A4 · SGLang dev-qwen38 |
262K | 2 |
| Tesla V100 x2 | 24.0 | Q8_K_XL · llama.cpp 9463 | 131K | 1 |
| Intel Arc Pro B70 | 27.8 | Q4_K_M · llama.cpp | 8K | 1 |
| Radeon AI Pro R9700 x2 | 43.6 | Q8_K_XL · llama.cpp | 262K | 1 |
Range: 16.8 → 435.1 tok/s. Nine engines. Twenty-five quant labels.
The same card, 26× apart
The RTX 5090's six rows are the clearest case:
| Format | Engine | Context | tok/s |
|---|---|---|---|
| IQ4_XS | lucebox |
131K | 435.1 |
| NVFP4 | TensorRT-LLM (NInfer) | 16K | 213.0 |
| W4A16 | vLLM cu129-nightly |
2K | 137.9 |
| W4A16 | vLLM v0.27.1 | 2K | 108.0 |
| UD-Q6_K | llama.cpp b10478 | 251K | 79.4 |
| FP8 | vLLM 0.27.1 | 8K | 16.8 |
A 26× spread on one card with one model. FP8 at 8K is the slowest row on the cheapest-to-reason-about configuration, and the same card's IQ4_XS row is the fastest in the entire 42-row set. Nothing about that is a property of the 5090.
The RX 7900 XTX shows the same effect at smaller scale: 154.2 tok/s with hipfire/MQ4 versus 58.0 with llama.cpp/Q4_K_M — 2.7× — on the same 24GB card, in the same month.
Practical read: if you are choosing a quantisation format and an engine for a 27B model, that choice moves throughput by more than upgrading the card does. Start there, then size the card for the context you need.
What a $1,300 used RTX 3090 actually delivers
The 3090 has the most rows (6) and the widest internal range of any card here:
| Format | Engine | Context | tok/s |
|---|---|---|---|
| EXL3 3.5bpw | exllamav3 1.4.2 (Mia fork) | 24K | 164.6 |
| EXL3 4.00bpw | exllamav2 v1.5.0 | 8K | 152.4 |
| EXL3 4.00bpw | buun-llama-cpp |
262K | 149.2 |
| Q4_K_XL | llama.cpp 0.3.0-dev | 24K | 81.9 |
| EXL3 4.00bpw | buun-llama-cpp |
150K | 27.8 |
| EXL3 4.00bpw | exllamav2 v1.5.0 | 150K | 24.5 |
Used 3090s were $1,300 in our 2026-09-28 snapshot (up from $1,050 on 2026-09-04); a new RTX 5090 had reached $9,000 (up from $4,899.99 on 2026-09-10, as US retail stock dried up). The 3090's best row is 164.6 tok/s versus 213.0 for the 5090's best non-exotic row (NVFP4 via TensorRT-LLM) — a 1.3× throughput gap against a 6.9× price gap. The 5090's advantage is not raw decode here; it is the 32GB of VRAM, the NVFP4 path, and the 435 tok/s outlier that a purpose-built engine reached.
The context tax is the real cliff
The steepest drop in this dataset is not a card difference at all. One submitter's 3090 runs of the same EXL3 4.00bpw build:
| Context | tok/s |
|---|---|
| 8K | 152.4 |
| 150K | 24.5 |
18× the context window for 6.2× less throughput — same GPU, same model, same quant, same engine version. Our earlier look at context scaling found that long context does not kill throughput on its own; what costs you is how the KV cache is quantised and how the engine schedules it. A 27B model at 128K+ is a memory-bandwidth problem, and the format you picked for the weights decides how much headroom is left.
Two caveats we are not going to paper over. First, a second submitter reported 149.2 tok/s at 262K on the same card and quant with a different llama.cpp fork — we cannot reconcile that with the 24.5 above, and both rows stay on the board with their source links. Second, these are 150K-row differences measured on rigs we did not control. Treat the shape as the finding, not the absolute numbers.
Old datacenter silicon is underrated
Three of the top ten rows are cards that cost a fraction of a modern consumer GPU, provided you can feed them:
- CMP 170HX — a repurposed mining card — reached 255.9 tok/s with W4A16 on vLLM 0.27.1 at 4.3K, and 140.5 tok/s with the same format at 65K. It also carries the only BF16 row in the set at 181.6 tok/s.
- Tesla V100 x4 hit 391.0 tok/s with NVFP4 on vLLM 1.2.2.
- Instinct MI100 x4 did 132.4 tok/s with INT8 at 64K.
The pattern is consistent: these parts have unusual memory configurations (HBM, 64GB+ per card) and the engine is doing the work. It also cuts the other way — a V100 x2 on Q8_K_XL managed 24.0 tok/s, slower than a single 3090, because the quant and engine mix did not suit the hardware.
What we would actually build for this model
| If you have | Run | Expected |
|---|---|---|
| One 24GB card (3090, $1,300 used) | EXL3 3.5–4bpw via exllamav3/exllamav2 | 150–165 tok/s at 8–24K |
| One 32GB card (5090, ~$9,000 new) | NVFP4 via TensorRT-LLM or SGLang | 195–213 tok/s, 16K+ |
| Two 24GB cards | AutoRound INT4 + FP8 KV via SGLang | ~220 tok/s at 262K |
| A workstation budget | RTX PRO 6000 Blackwell, 96GB | 296 tok/s at 262K |
A 48GB RTX A6000 was $5,975.06 new in the 2026-09-10 snapshot — more than two 3090s, for less throughput than two 3090s delivered on this model. VRAM is worth paying for; VRAM density at that price is not.
The board's Qwen3.8-27B profile has all 42 rows with their per-row context and source, and the GPU comparison puts two cards side by side per model. If you have a run we do not, submit it — the numbers here came from people who did.
FAQ
How much VRAM does Qwen3.8-27B need?
A 27–28B model fits in 24GB at 4-bit quantisation with room for a modest context window; our 3090 rows at 24K context are the proof. Above ~100K context, KV cache becomes the constraint and the same card drops from ~152 to ~25 tok/s.
Which GPU is fastest for Qwen3.8-27B?
The fastest verified row is an RTX 5090 at 435.1 tok/s with IQ4_XS on a custom engine, followed by a Tesla V100 x4 at 391.0 with NVFP4. Both results come from a specific format-plus-engine pairing, not from the card alone — the same 5090 runs the model at 16.8 tok/s with FP8 under vLLM.
Is quantisation or the GPU more important for throughput?
On this dataset, quantisation and engine choice move throughput more than a tier of GPU does: 16.8 → 435.1 tok/s on one RTX 5090, versus 164.6 → 213.0 tok/s between the best 3090 and the best non-exotic 5090 row.
How do you verify these numbers?
Every row is imported from a source link and carries hardware, quant, engine version and context — the four items our methodology treats as the minimum for a benchmark to mean anything. Rows missing any of them stay out of the board.
Sources
- Board rows, Qwen3.8-27B: model profile (42 verified rows, pulled 2026-09-28).
- RTX 4090 int4 output-head-only / vLLM 141.2 @38K — x.com/yume_arasaki/status/2099649396917121109.
- RTX 5090 IQ4_XS /
lucebox435.1 @131K — localmaxxing runcmtldroid00iwoe017o075049. - RTX 5090 FP8 / vLLM 0.27.1 16.8 @8K — localmaxxing run
cmstify5k03m4ms01dz1byloh. - RTX 5090 W4A16 / vLLM 108.0 and 137.9 @2K — localmaxxing runs
cmsxhsbwo0bazms01nkujyy99,cmt4ektr900edpu01ozgyp05z. - RTX 3090 EXL3 4.00bpw 152.4 @8K and 24.5 @150K — x.com/mr_r0b0t/status/2100226555888734375; 149.2 @262K and 27.8 @150K — x.com/mr_r0b0t/status/2101858002051412021.
- RTX 3090 EXL3 3.5bpw 164.6 @24K — localmaxxing run
cmtmoa067001alt01win7qp66. - CMP 170HX W4A16 255.9 @4.3K — localmaxxing run
cmtkw0izf042hp701ucbvsc3l; BF16 181.6 @4.3K — runcmtkowbb903v0p701m945wvc9; INT8 147.7 and W4A16 140.5 @65K — x.com/seanphan/status/2094258885146370385. - Prices: used RTX 3090 $1,300 (2026-09-28 snapshot; $1,050 on 2026-09-04); new RTX 5090 ~$9,000 (2026-09-28 snapshot; $4,899.99 on 2026-09-10) — see the September wrap for the move; new RTX A6000 48GB $5,975.06 (Newegg, 2026-09-10); new Radeon AI Pro R9700 $1,329 (Newegg, 2026-09-04); Intel Arc Pro B70 $949.99 (Newegg, 2026-09-28). See the price reports.