DiggTop ⌕

M4 Max 128GB vs RTX 5090

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

M4 Max 128GB1 verified run · best 134.0 tok/s
RTX 509024 verified runs · best 2231.4 tok/s
ModelM4 Max 128GBRTX 5090
LFM2-350M—2231.4 NVFP4
NVIDIA-Nemotron-3.5-Lightning-30B-A3B—446.7 NVFP4
Qwen3.8-27B134.0 4-bit435.1 IQ4_XS
Ornith-1.0-35B-PrismaAURA-4.75bit-vllm-MTP—283.1 NVFP4
Kimi-Linear-48B-A3B-Instruct—280.8 IQ4_XS
Ornith-1.5-35B-A3B—258.8 NVFP4
North-Mini-Code-1.0—258.4 UD-Q6_K_XL
Qwen3.6-35B-A3B-MTP—222.9 Unsloth-Dynamic-IQ4_XS
Laguna-XS-2.1—216.7 NVFP4
KAT-Coder-V2.5-Dev-AWQ-W4A16-ASYM—201.2 W4A16
Ornith-1.0-35B—196.6 AWQ
Muse-Glimmer-30B—124.8 NVFP4
Qwen3.6-27B—120.9 NVFP4
Huihui-Qwen3.8-27B-abliterated—117.7 NVFP4
Qwen3.8-27B-5090-goldilocks—93.5 Q6_K
Qwen3.6-27B-MTP—28.3 Q5_K_M
DeepSeek-V4-Flash-0731-MXFP4—26.4 MXFP4

Head-to-head (models tested on both): M4 Max 128GB 0 · RTX 5090 1 · tied 0 · data as of Sep 29, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board

↑