DiggTop

Multi-GPU x3 vs RTX PRO 6000 Blackwell

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

Multi-GPU x34 verified runs · best 63.3 tok/s
RTX PRO 6000 Blackwell29 verified runs · best 1271.3 tok/s
ModelMulti-GPU x3RTX PRO 6000 Blackwell
MobileLLM-125M1271.3 Q8_0
Qwen2.5-0.5B-Instruct1060.0 Q8_0
MobileLLM-350M980.8 Q8_0
SmolLM2-360M-Instruct968.1 Q8_0
Qwen3-0.6B956.4 Q4_K_M
Llama-3.2-1B-Instruct699.9 Q8_0
gemma-3-1b-it664.3 Q4_K_M
Llama-3.2-3B-Instruct489.2 Q4_K_M
Ling-3.0-tiny369.1 Q4_K_M
Qwen3.6-35B-A3B63.3 APEX-I-Balanced330.2 NVFP4
Ornith-1.0-35B309.7 Q4_K_M
Qwen3.8-27B296.0 NVFP4
DeepSeek-Coder-V2-Lite-Instruct237.3 FP8
Qwen3-Coder-Next174.3 FP8
Llama-3.1-8B-Instruct158.1 FP8
Muse-Glimmer-30B152.2 NVFP4
Qwen3.8-27B-MTP112.1 NVFP4
Qwen3.6-27B60.3 Q6_K_L101.4 FP8
Qwen3-14B90.8 FP8
Kimi-Linear-48B-A3B-Instruct90.0 Q8_0
Qwen3.6-27B-MTP60.1 Q6_K
Qwen2.5-Coder-14B-Instruct50.3 BF16
Ling-3.0-flash43.4 AD-IQ4_XXS
Qwen3-Coder-30B-A3B-Instruct26.1 FP8
Laguna-S-2.123.9 UD-IQ4_XS

Head-to-head (models tested on both): Multi-GPU x3 0 · RTX PRO 6000 Blackwell 2 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board