DiggTop

CMP 170HX vs RTX PRO 6000 Blackwell

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

CMP 170HX16 verified runs · best 255.9 tok/s
RTX PRO 6000 Blackwell29 verified runs · best 1271.3 tok/s
ModelCMP 170HXRTX PRO 6000 Blackwell
MobileLLM-125M1271.3 Q8_0
Qwen2.5-0.5B-Instruct1060.0 Q8_0
MobileLLM-350M980.8 Q8_0
SmolLM2-360M-Instruct968.1 Q8_0
Qwen3-0.6B956.4 Q4_K_M
Llama-3.2-1B-Instruct699.9 Q8_0
gemma-3-1b-it664.3 Q4_K_M
Llama-3.2-3B-Instruct489.2 Q4_K_M
Ling-3.0-tiny369.1 Q4_K_M
Qwen3.6-35B-A3B87.2 UD-Q8_K_XL330.2 NVFP4
Ornith-1.0-35B309.7 Q4_K_M
Qwen3.8-27B255.9 W4A16296.0 NVFP4
DeepSeek-Coder-V2-Lite-Instruct237.3 FP8
Qwen3-Coder-Next174.3 FP8
NVIDIA-Nemotron-3.5-Lightning-30B-A3B165.1 Q8_0
Llama-3.1-8B-Instruct158.1 FP8
Muse-Glimmer-30B152.2 NVFP4
Qwen3.8-27B-MTP46.4 INT8-W8A16112.1 NVFP4
Qwen3.6-27B101.4 FP8
Qwen3-14B90.8 FP8
Kimi-Linear-48B-A3B-Instruct90.0 Q8_0
Ornith-1.5-35B-A3B87.1 Q8_0
Ornith-1.5-9B61.5 BF16
Qwen3.6-27B-MTP60.1 Q6_K
Qwen2.5-Coder-14B-Instruct50.3 BF16
Qwen3.8-Flash-Next46.1 UD-IQ1_S
Llama-4-Scout-17B-16E-Instruct44.5 UD-Q3_K_XL
gemma-4-31B-it28.3 NVFP4
Qwen3-Coder-30B-A3B-Instruct26.1 FP8
granite-4.2-30b23.7 Q8_0
Llama-3.3-70B-Instruct20.1 NVFP4

Head-to-head (models tested on both): CMP 170HX 0 · RTX PRO 6000 Blackwell 3 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board