DiggTop

Multi-GPU x3 vs RTX 3060 x2

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

Multi-GPU x34 verified runs · best 63.3 tok/s
RTX 3060 x229 verified runs · best 198.1 tok/s
ModelMulti-GPU x3RTX 3060 x2
Qwen3-1.7B198.1 Q4_K_M
MiniCPM5-1B125.9 Q4_K_M
Trinity-Mini109.0 Q4_K_M
Phi-4-mini-reasoning108.2 Q4_K_M
gemma-4-26B-A4B-it-qat95.3 Q4_K_M
Qwen3-30B-A3B-Thinking-250794.3 Q4_K_M
gpt-oss-20b93.4 Q4_K_M
OpenReasoning-Nemotron-7B68.0 Q4_K_M
Qwen3.6-35B-A3B63.3 APEX-I-Balanced
WebWorld-8B62.3 Q4_K_M
gemma-4-26B-A4B-it62.1 UD-Q4_K_M
Qwen3.6-27B60.3 Q6_K_L18.9 IQ4_NL
Carnice-9b56.1 Q4_K_M
Muse-Glimmer-30B56.0 UD-Q4_K_XL
Ling-3.0-flash43.4 AD-IQ4_XXS
NVIDIA-Nemotron-Nano-12B-v240.4 Q4_K_M
gemma-3-12b-it39.4 Q4_K_M
Qwen3-14B35.8 Q4_K_M
Nemotron-Cascade-14B-Thinking35.7 Q4_K_M
Phi-4-reasoning35.4 Q4_K_M
WebWorld-14B35.4 Q4_K_M
InternVL3-14B35.1 Q4_K_M
DeepSeek-R1-Distill-Qwen-14B35.1 Q4_K_M
phi-434.2 Q4_K_M
DeepHermes-3-Mistral-24B-Preview24.5 Q4_K_M
reka-flash-323.9 Q4_K_M
Laguna-S-2.123.9 UD-IQ4_XS
mistralai_Mistral-Small-3.2-24B-Instruct-250621.8 Q4_K_M
Tess-4-27B18.8 Q4_K_M
GLM-Z1-32B-041416.6 Q4_K_M
Olmo-3-32B-Think16.4 Q4_K_M
gemma-4-31B-it15.8 Q4_K_M

Head-to-head (models tested on both): Multi-GPU x3 1 · RTX 3060 x2 0 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board