DiggTop ⌕

M4 Max 128GB vs RTX PRO 6000 Blackwell

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

M4 Max 128GB1 verified run · best 134.0 tok/s
RTX PRO 6000 Blackwell29 verified runs · best 1271.3 tok/s
ModelM4 Max 128GBRTX PRO 6000 Blackwell
MobileLLM-125M—1271.3 Q8_0
Qwen2.5-0.5B-Instruct—1060.0 Q8_0
MobileLLM-350M—980.8 Q8_0
SmolLM2-360M-Instruct—968.1 Q8_0
Qwen3-0.6B—956.4 Q4_K_M
Llama-3.2-1B-Instruct—699.9 Q8_0
gemma-3-1b-it—664.3 Q4_K_M
Llama-3.2-3B-Instruct—489.2 Q4_K_M
Ling-3.0-tiny—369.1 Q4_K_M
Qwen3.6-35B-A3B—330.2 NVFP4
Ornith-1.0-35B—309.7 Q4_K_M
Qwen3.8-27B134.0 4-bit296.0 NVFP4
DeepSeek-Coder-V2-Lite-Instruct—237.3 FP8
Qwen3-Coder-Next—174.3 FP8
Llama-3.1-8B-Instruct—158.1 FP8
Muse-Glimmer-30B—152.2 NVFP4
Qwen3.8-27B-MTP—112.1 NVFP4
Qwen3.6-27B—101.4 FP8
Qwen3-14B—90.8 FP8
Kimi-Linear-48B-A3B-Instruct—90.0 Q8_0
Qwen3.6-27B-MTP—60.1 Q6_K
Qwen2.5-Coder-14B-Instruct—50.3 BF16
Qwen3-Coder-30B-A3B-Instruct—26.1 FP8

Head-to-head (models tested on both): M4 Max 128GB 0 · RTX PRO 6000 Blackwell 1 · tied 0 · data as of Sep 29, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board

↑