DiggTop ⌕

Mac Studio (Mac16,9) with Apple M4 Max vs RTX 3090

Best sourced output tok/s per model (single-stream generation) — engine and version, quantization, context and a source link on every row. Hover a value for the quant.

Mac Studio (Mac16,9) with Apple M4 Max1 sourced run · best 110.3 tok/s
RTX 309018 verified runs · best 709.5 tok/s
ModelMac Studio (Mac16,9) with Apple M4 MaxRTX 3090
LFM2.5-1.2B-Instruct—709.5 Q4_K_M
Qwen3.5-0.8B-MTP—476.1 Q6_K_XL
NVIDIA-Nemotron-3.5-Lightning-30B-A3B—353.7 NVFP4
LFM2.5-2.6B—305.6 Q4_K_M
Qwen3.5-4B-MTP—240.4 Q6_K_XL
Qwen3.8-27B—164.6 EXL3-3.5bpw
Qwen3.5-9B-MTP—164.3 Q6_K_XL
Qwen3.8-Flash-Next110.3 IQ3_XXS + MTP Q8_038.6 EXL3 2.50bpw
Ornstein3.6-27B-MTP-NSC-ACE-SABER—32.9 Q4_K_M
GLM-5.3-Flash—28.1 EXL3 3.05 bpw
DeepSeek-V4.1-Flash—26.0 EXL3 3.0 bpw

Head-to-head (models tested on both): Mac Studio (Mac16,9) with Apple M4 Max 1 · RTX 3090 0 · tied 0 · data as of Oct 11, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board

↑