DiggTop ⌕

DGX Spark vs Mac Studio (Mac16,9) with Apple M4 Max

Best sourced output tok/s per model (single-stream generation) — engine and version, quantization, context and a source link on every row. Hover a value for the quant.

DGX Spark47 sourced runs · best 86.8 tok/s
Mac Studio (Mac16,9) with Apple M4 Max1 verified run · best 110.3 tok/s
ModelDGX SparkMac Studio (Mac16,9) with Apple M4 Max
Qwen3.8-Flash-Next75.7 4.05bpw_h6_ng6110.3 IQ3_XXS + MTP Q8_0
Qwen 3.8 Flash86.8 EXL3 3.05bpw—
GLM-5.3-Flash64.1 EXL3 2.05 bpw—
CYBER-FROST-3.860.7 EXL3—
Qwen3.8-Flash-Next-hibrid4860.0 NVFP4 (output head)—
Qwen3.8-27B56.0 NVFP4 W4A4—
DeepSeek-V4.1-Flash47.0 FP4—
Underdog-Saluki-27B22.6 IQ2-mix—

Head-to-head (models tested on both): DGX Spark 0 · Mac Studio (Mac16,9) with Apple M4 Max 1 · tied 0 · data as of Oct 11, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board

↑