CMP 170HX, B70, V100: Cheap Cards Punching Up
A new RTX 5090 was $9,000 in our 2026-09-28 snapshot. A used RTX 3090 was $1,300, a used CMP 170HX $1,050, a new Intel Arc Pro B70 $949.99. On the one model all of them have actually run, the cheapest card in that list owns the fastest matched row.
It is not one weird benchmark. It is what happens when you compare workloads instead of cards: the 64GB HBM2e mining card, the Intel SYCL card, four Volta parts and an AMD gfx908 stack each win somewhere and lose badly elsewhere.
As of 2026-09-28. Every figure below is an imported row from an upstream source that we verified — we did not run these tests. Prices are weekly snapshot items; see prices and the sources at the end.
The one model all of them ran
Qwen3.8-27B is the only model with verified rows on the CMP 170HX, the Arc Pro B70, the RTX 3090, the RTX 5090, a Tesla V100 x4 and an Instinct MI100 x4. Best verified row per card:
| Card | Best row | Format · engine | Context | Rows |
|---|---|---|---|---|
| RTX 5090 ($9,000) | 435.1 | IQ4_XS · lucebox |
131K | 24 |
| Tesla V100 x4 (no price) | 391.0 | NVFP4 · vLLM 1.2.2 | 32K | 1 |
| CMP 170HX ($1,050) | 255.9 | W4A16 · vLLM 0.27.1 | 4.3K | 16 |
| RTX 3090 ($1,300) | 164.6 | EXL3 3.5bpw · exllamav3 1.4.2 | 24K | 16 |
| Instinct MI100 x4 (no price) | 132.4 | INT8 · vLLM (gfx908 fork) | 64K | 1 |
| Intel Arc Pro B70 x2 (~$1,900) | 101.9 | AutoRound INT4 W4A16 · vLLM | 2K | 5 |
| Intel Arc Pro B70 ($949.99) | 27.8 | Q4_K_M · llama.cpp | 8K | 16 |
Read the left column as best available format, not card speed: the 5090's 435.1 and the 3090's 164.6 come from different engines on different quants, so that pair is indicative, not measured. Per-row detail is on each card page and on benchmarks.
The matched comparison: W4A16 on vLLM
Three cards ran the same model, same quantisation class and same engine family — the closest thing to a controlled test here:
| Card | Quant · engine | Context | tok/s |
|---|---|---|---|
| CMP 170HX — $1,050 | W4A16 · vLLM 0.27.1 | 4.3K | 255.9 |
| RTX 5090 — $9,000 | W4A16 · vLLM cu129-nightly |
2K | 137.9 |
| Intel Arc Pro B70 x2 — ~$1,900 | AutoRound INT4 W4A16 · vLLM | 2K | 101.9 |
The $1,050 card is 1.86× faster than the card that costs 8.6× more, same model, same weight format. One row down: the 5090's W4A16 run on vLLM 0.27.1 is 108.0 tok/s, and the CMP's INT8 row at 65K context is 147.7 — faster than the 5090's best W4A16 row at two thousand tokens of context. See CMP 170HX vs RTX 5090.
What the 5090 buys at that price is not dense-27B decode throughput. It is 32GB of GDDR7, an NVFP4 path at 213.0 tok/s via TensorRT-LLM, and a 435.1 tok/s purpose-built engine. Those matter — just not in the way "7× the price" implies.
Four MI100s lose to one mining card
The most useful matched pair in the set is INT8 at ~64K context:
| Card | Quant · engine | Context | tok/s | Session VRAM |
|---|---|---|---|---|
| CMP 170HX (one card) | INT8 · vLLM 0.27.1 | 65K | 147.7 | 64GB |
| Instinct MI100 x4 | INT8 · vLLM mi100-optimized-sy fork |
64K | 132.4 | 128GB |
One used card beat four AMD datacenter cards on the same model, quantisation and a context window within 1,500 tokens. That MI100 row is a single submitted row on a community-forked vLLM for gfx908, so treat the absolute number as fragile. We have no current price for the MI100 — or the V100, or the RTX PRO 6000 — so this is a throughput comparison only, with no per-dollar claim attached.
Where the cheap tier disappoints
- The Arc Pro B70 is not a dense-27B card. Qwen3.8-27B at Q4_K_M via llama.cpp is 27.8 tok/s on one card, 101.9 on a two-card vLLM setup. An RTX 3090 does 81.9 on the same model with the similar Q4-class Q4_K_XL quant at 24K context — 2.95× the single-B70 figure. The B70's real strength is MoE and long context: 283.1 tok/s on Laguna-XS-2.1 (33B, Q5_K_M) at 131K, 248.3 on Qwen3.6-35B-A3B, 224.9 on Ornith-1.0-35B-B70-Turbo. All three are different models from the 5090's 435.1, so cross-model rankings here are indicative only.
- Tesla V100 x2 falls off a cliff at long context. Qwen3.8-27B at Q8_K_XL and 131K is 24.0 tok/s on two V100s — under a single 3090's 24.5 at 150K. One V100 on Qwen3.6-35B-A3B at 262K managed 4.3 tok/s. The one hero row is x4 at 391.0 on NVFP4, and NVFP4 is not a format Volta supports natively; we publish it with its source but do not treat it as a hardware capability. Our own hardware table also lists 16GB and 32GB for different V100 entries, so V100 VRAM sizing here is unverified.
- The CMP 170HX has hard limits. Llama-3.3-70B at NVFP4 and 32K is 20.1 tok/s using 57GB of VRAM. Ornith-1.5-35B-A3B at 131K drops to 87.1 (Q8_0) and 48.6 (W4A16), and ten stacked cards ran GLM-5.2 at 29.7. Our hardware table records 64GB, and the price source describes a community memory-unlock mod — so treat the 64GB figure as a modified-card claim, not a factory spec.
- The 5090 is not immune. Its worst Qwen3.8-27B row is 16.8 tok/s (FP8, vLLM 0.27.1, 8K), and a neighbouring 27B model at 131K through llama.cpp logged 4.1 tok/s. At $9,000, per-dollar throughput is the weakest column in the table.
What the prices actually say
Tok/s per $100 of card price, from the rows above. The two-card B70 figure assumes a second card at the same $949.99 unit price.
| Card | Price (2026-09-28) | Matched row | tok/s per $100 |
|---|---|---|---|
| Intel Arc Pro B70 (MoE, 33B) | $949.99 new | 283.1 @131K | 29.80 |
| CMP 170HX (W4A16, vLLM) | $1,050 used | 255.9 @4.3K | 24.37 |
| RTX 3090 (EXL3 3.5bpw) | $1,300 used | 164.6 @24K | 12.66 |
| Arc Pro B70 x2 (INT4, vLLM) | ~$1,899.98 | 101.9 @2K | 5.36 |
| RTX 5090 (best row) | $9,000 new | 435.1 @131K | 4.83 |
| Arc Pro B70 (dense 27B) | $949.99 new | 27.8 @8K | 2.93 |
| RTX 5090 (W4A16, vLLM) | $9,000 new | 137.9 @2K | 1.53 |
The CMP pays off on a matched dense-27B workload; the Arc Pro B70 pays off on MoE, where it is the cheapest per unit of throughput in the set — provided your model is the kind it is good at. An RTX 5060 Ti 16GB was $800 and an RTX 3060 $495 in the same snapshot; neither appears in the Qwen3.8-27B verified set. Compare cards per model with compare and check value before you bid.
What these cards cannot do
- The software stack is the price of admission. The Arc Pro B70 rows run on oneAPI 2026.1.1 SYCL builds, custom XPU vLLM trees and campaign-specific kernels; the MI100 row runs a community
gfx908fork. You are buying a build recipe, not a supported product. - Driver, platform and cooling friction. The CMP 170HX has no display outputs — it is a second card by design. V100-era parts want server airflow and often modified power cabling. These are 250–300W-class parts; a passive card on a desk is a thermal problem you must solve.
- No warranty, no resale floor. A memory-unlocked mining card is fully depreciated the day it arrives. If it dies, the run history on this page does not help you.
- Some rows are load-bearing singletons. The V100 x4 (391.0) and MI100 x4 (132.4) headlines are one verified row each from one submitter. Our methodology requires engine, quant and context before a row counts, but it cannot conjure a second independent run.
If you own one of these and have a run we lack, submit it — the matched W4A16 pair above exists because two different people published theirs.
FAQ
What is the best GPU for local LLM inference per dollar?
In our 2026-09-28 data, a used CMP 170HX at $1,050 returned 255.9 tok/s on Qwen3.8-27B with W4A16 on vLLM — 24.4 tok/s per $100, against 4.8 tok/s per $100 for a $9,000 RTX 5090 on its best row. A used RTX 3090 at $1,300 is next on that measure at 12.7.
Is the CMP 170HX worth buying for LLM inference?
On short-context dense 27–35B work it is the strongest price/performance card we track. At 32K on Llama-3.3-70B it falls to 20.1 tok/s, and at 131K on a 35B MoE to 87.1. Add no display outputs, a community 64GB memory unlock, no warranty and driver fiddling.
Can a Tesla V100 or Instinct MI100 x4 keep up with a modern consumer GPU?
On the same model, no. Four MI100s reached 132.4 tok/s on Qwen3.8-27B INT8 at 64K, below a single CMP 170HX's 147.7 at 65K. A V100 x4 posted 391.0 on an NVFP4 row we treat as indicative, while a V100 x2 managed only 24.0 at 131K. Neither card has a price row in our snapshot, so we compare throughput only.
Why is a $9,000 RTX 5090 slower than a $1,050 CMP 170HX?
On the matched W4A16/vLLM comparison it is: 137.9 tok/s at 2K versus 255.9 at 4.3K. The 5090's edge is elsewhere — 32GB of GDDR7, NVFP4 via TensorRT-LLM at 213.0, and one 435.1 row from a purpose-built engine at 131K.
Are these benchmark numbers yours?
No. Every row is imported from an upstream source, carries hardware, quant, engine version and context, and was verified before it entered the board — but we did not run these tests. We name the engine and context next to every number for that reason.
Sources
- Per-card verified coverage and all rows: benchmarks (pulled 2026-09-28). Card pages: CMP 170HX · CMP 170HX x2 · Intel Arc Pro B70 · Arc Pro B70 x2 · RTX 3090 · RTX 5090 · Tesla V100 x4 · Instinct MI100 x4.
- CMP 170HX Qwen3.8-27B W4A16 255.9 @4.3K, BF16 181.6 — localmaxxing
cmtkw0izf042hp701ucbvsc3l,cmtkowbb903v0p701m945wvc9; INT8 147.7 and W4A16 140.5 @65K — x.com/seanphan/status/2094258885146370385; Llama-3.3-70B 20.1 —cmtkzgp1t0066oe01mab72fjd; Ornith-1.5-35B-A3B 87.1/48.6 @131K —cmtl0zx5v0087oe0188neqqjc,cmtkumtz1041mp701czybzv0j; GLM-5.2 x10 29.7 —cmtlp82z000rhoe01qmrvpq9i. - RTX 5090 Qwen3.8-27B IQ4_XS 435.1 —
cmtldroid00iwoe017o075049; NVFP4 213.0 —cmtkgxt0003rqp7019bcxtyyn; W4A16 137.9 / 108.0 —cmt4ektr900edpu01ozgyp05z,cmsxhsbwo0bazms01nkujyy99; FP8 16.8 —cmstify5k03m4ms01dz1byloh; UD-Q6_K 79.4 —cmtllqn4800naoe01pngxlyjr; Qwen3.6-27B 4.1 —cmstyd8wp0487ms01jlrmk845. - RTX 3090 Qwen3.8-27B EXL3 3.5bpw 164.6 —
cmtmoa067001alt01win7qp66; EXL3 4.00bpw 152.4/24.5 @8K/150K — x.com/mr_r0b0t/status/2100226555888734375; 149.2/27.8 @262K/150K — x.com/mr_r0b0t/status/2101858002051412021; Q4_K_XL 81.9 —cmtmo9u6p000vlt01sjmhhfm9; Nemotron-3.5-Lightning NVFP4 353.7 —cmspe7wa500pwmp01c2szra70. - Arc Pro B70 Laguna-XS-2.1 283.1, Ornith-1.0-35B 224.9, Qwen3.6-35B-A3B 248.3, Qwen3.8-27B 27.8, Muse-Glimmer-30B 29.2 — localmaxxing
cmr89jid800a9qr01lkojydrn,cmr5x1qce00khq901hg2qfi09,cmsk153ne00q1qm01369d61m8,cmt9m8i0b00z3li01o1ragvte,cmsp799tx00mnmp016yziebh4. Arc Pro B70 x2 Qwen3.8-27B 101.9/67.7/136.4, Flash-Next 20.3 —cmszbkxco0e11ms01l2rixxbt,cmstyaenj046xms01azvmwlpo,cmt0vu76q0fvtms017exhssx9,cmtigo39n04e2p401eumx5joo. - Tesla V100 x4 NVFP4 391.0 @32K —
cmt1l1boi02kxmv010f61ghs2; V100 x2 24.0 @131K and Ornith-1.5-35B-A3B 65.0 —cmthoeuma0247p401x27r5r1b,cmt2qvl780gwqmv01po181spa, Flash-Next 96.1/87.5/80.9 — x.com/PeasantSmith/status/2102804368286175514; V100 Qwen3.6-35B-A3B 4.3 @262K —cmsgihjlw017qpp01lib71256. - Instinct MI100 x4 INT8 132.4 @64K, 128GB peak —
cmt7l63b7003ynn01s6jp12vs. - Prices, snapshot 2026-09-28: RTX 5090 $9,000 new (
us-street-avg, ixbt); used RTX 3090 $1,300 (gamebastion); CMP 170HX $1,050 used (overclockers.ua); Intel Arc Pro B70 $949.99 new (TechPowerUp); RTX 5060 Ti 16GB $800 (wccftech); RTX 3060 $495 (wccftech). No 2026-09-28 price row exists for Tesla V100, Instinct MI100 or RTX PRO 6000. All price reports.
Prices are weekly-snapshot items. Re-check prices before buying.