DiggTop ⌕

CMP 170HX, B70, V100: Cheap Cards Punching Up

A new RTX 5090 was $9,000 in our 2026-09-28 snapshot. A used RTX 3090 was $1,300, a used CMP 170HX $1,050, a new Intel Arc Pro B70 $949.99. On the one model all of them have actually run, the cheapest card in that list owns the fastest matched row.

It is not one weird benchmark. It is what happens when you compare workloads instead of cards: the 64GB HBM2e mining card, the Intel SYCL card, four Volta parts and an AMD gfx908 stack each win somewhere and lose badly elsewhere.

As of 2026-09-28. Every figure below is an imported row from an upstream source that we verified — we did not run these tests. Prices are weekly snapshot items; see prices and the sources at the end.

The one model all of them ran

Qwen3.8-27B is the only model with verified rows on the CMP 170HX, the Arc Pro B70, the RTX 3090, the RTX 5090, a Tesla V100 x4 and an Instinct MI100 x4. Best verified row per card:

Card Best row Format · engine Context Rows
RTX 5090 ($9,000) 435.1 IQ4_XS · lucebox 131K 24
Tesla V100 x4 (no price) 391.0 NVFP4 · vLLM 1.2.2 32K 1
CMP 170HX ($1,050) 255.9 W4A16 · vLLM 0.27.1 4.3K 16
RTX 3090 ($1,300) 164.6 EXL3 3.5bpw · exllamav3 1.4.2 24K 16
Instinct MI100 x4 (no price) 132.4 INT8 · vLLM (gfx908 fork) 64K 1
Intel Arc Pro B70 x2 (~$1,900) 101.9 AutoRound INT4 W4A16 · vLLM 2K 5
Intel Arc Pro B70 ($949.99) 27.8 Q4_K_M · llama.cpp 8K 16

Read the left column as best available format, not card speed: the 5090's 435.1 and the 3090's 164.6 come from different engines on different quants, so that pair is indicative, not measured. Per-row detail is on each card page and on benchmarks.

The matched comparison: W4A16 on vLLM

Three cards ran the same model, same quantisation class and same engine family — the closest thing to a controlled test here:

Card Quant · engine Context tok/s
CMP 170HX — $1,050 W4A16 · vLLM 0.27.1 4.3K 255.9
RTX 5090 — $9,000 W4A16 · vLLM cu129-nightly 2K 137.9
Intel Arc Pro B70 x2 — ~$1,900 AutoRound INT4 W4A16 · vLLM 2K 101.9

The $1,050 card is 1.86× faster than the card that costs 8.6× more, same model, same weight format. One row down: the 5090's W4A16 run on vLLM 0.27.1 is 108.0 tok/s, and the CMP's INT8 row at 65K context is 147.7 — faster than the 5090's best W4A16 row at two thousand tokens of context. See CMP 170HX vs RTX 5090.

What the 5090 buys at that price is not dense-27B decode throughput. It is 32GB of GDDR7, an NVFP4 path at 213.0 tok/s via TensorRT-LLM, and a 435.1 tok/s purpose-built engine. Those matter — just not in the way "7× the price" implies.

Four MI100s lose to one mining card

The most useful matched pair in the set is INT8 at ~64K context:

Card Quant · engine Context tok/s Session VRAM
CMP 170HX (one card) INT8 · vLLM 0.27.1 65K 147.7 64GB
Instinct MI100 x4 INT8 · vLLM mi100-optimized-sy fork 64K 132.4 128GB

One used card beat four AMD datacenter cards on the same model, quantisation and a context window within 1,500 tokens. That MI100 row is a single submitted row on a community-forked vLLM for gfx908, so treat the absolute number as fragile. We have no current price for the MI100 — or the V100, or the RTX PRO 6000 — so this is a throughput comparison only, with no per-dollar claim attached.

Where the cheap tier disappoints

What the prices actually say

Tok/s per $100 of card price, from the rows above. The two-card B70 figure assumes a second card at the same $949.99 unit price.

Card Price (2026-09-28) Matched row tok/s per $100
Intel Arc Pro B70 (MoE, 33B) $949.99 new 283.1 @131K 29.80
CMP 170HX (W4A16, vLLM) $1,050 used 255.9 @4.3K 24.37
RTX 3090 (EXL3 3.5bpw) $1,300 used 164.6 @24K 12.66
Arc Pro B70 x2 (INT4, vLLM) ~$1,899.98 101.9 @2K 5.36
RTX 5090 (best row) $9,000 new 435.1 @131K 4.83
Arc Pro B70 (dense 27B) $949.99 new 27.8 @8K 2.93
RTX 5090 (W4A16, vLLM) $9,000 new 137.9 @2K 1.53

The CMP pays off on a matched dense-27B workload; the Arc Pro B70 pays off on MoE, where it is the cheapest per unit of throughput in the set — provided your model is the kind it is good at. An RTX 5060 Ti 16GB was $800 and an RTX 3060 $495 in the same snapshot; neither appears in the Qwen3.8-27B verified set. Compare cards per model with compare and check value before you bid.

What these cards cannot do

If you own one of these and have a run we lack, submit it — the matched W4A16 pair above exists because two different people published theirs.

FAQ

What is the best GPU for local LLM inference per dollar?

In our 2026-09-28 data, a used CMP 170HX at $1,050 returned 255.9 tok/s on Qwen3.8-27B with W4A16 on vLLM — 24.4 tok/s per $100, against 4.8 tok/s per $100 for a $9,000 RTX 5090 on its best row. A used RTX 3090 at $1,300 is next on that measure at 12.7.

Is the CMP 170HX worth buying for LLM inference?

On short-context dense 27–35B work it is the strongest price/performance card we track. At 32K on Llama-3.3-70B it falls to 20.1 tok/s, and at 131K on a 35B MoE to 87.1. Add no display outputs, a community 64GB memory unlock, no warranty and driver fiddling.

Can a Tesla V100 or Instinct MI100 x4 keep up with a modern consumer GPU?

On the same model, no. Four MI100s reached 132.4 tok/s on Qwen3.8-27B INT8 at 64K, below a single CMP 170HX's 147.7 at 65K. A V100 x4 posted 391.0 on an NVFP4 row we treat as indicative, while a V100 x2 managed only 24.0 at 131K. Neither card has a price row in our snapshot, so we compare throughput only.

Why is a $9,000 RTX 5090 slower than a $1,050 CMP 170HX?

On the matched W4A16/vLLM comparison it is: 137.9 tok/s at 2K versus 255.9 at 4.3K. The 5090's edge is elsewhere — 32GB of GDDR7, NVFP4 via TensorRT-LLM at 213.0, and one 435.1 row from a purpose-built engine at 131K.

Are these benchmark numbers yours?

No. Every row is imported from an upstream source, carries hardware, quant, engine version and context, and was verified before it entered the board — but we did not run these tests. We name the engine and context next to every number for that reason.

Sources

Prices are weekly-snapshot items. Re-check prices before buying.

↑