DiggTop

CMP 170HX vs Multi-GPU x3

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

CMP 170HX16 verified runs · best 255.9 tok/s
Multi-GPU x34 verified runs · best 63.3 tok/s
ModelCMP 170HXMulti-GPU x3
Qwen3.8-27B255.9 W4A16
NVIDIA-Nemotron-3.5-Lightning-30B-A3B165.1 Q8_0
Qwen3.6-35B-A3B87.2 UD-Q8_K_XL63.3 APEX-I-Balanced
Ornith-1.5-35B-A3B87.1 Q8_0
Ornith-1.5-9B61.5 BF16
Qwen3.6-27B60.3 Q6_K_L
Qwen3.8-27B-MTP46.4 INT8-W8A16
Qwen3.8-Flash-Next46.1 UD-IQ1_S
Llama-4-Scout-17B-16E-Instruct44.5 UD-Q3_K_XL
Ling-3.0-flash43.4 AD-IQ4_XXS
gemma-4-31B-it28.3 NVFP4
Laguna-S-2.123.9 UD-IQ4_XS
granite-4.2-30b23.7 Q8_0
Llama-3.3-70B-Instruct20.1 NVFP4

Head-to-head (models tested on both): CMP 170HX 1 · Multi-GPU x3 0 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board