RTX 3060 x2 vs RTX 3090 x2
Best verified output tok/s per model (single-stream generation). Hover a value for the quant.
RTX 3060 x229 verified runs · best 198.1 tok/s
RTX 3090 x26 verified runs · best 240.8 tok/s
| Model | RTX 3060 x2 | RTX 3090 x2 |
|---|---|---|
| Qwen3.6-27B | 18.9 | 240.8 |
| Qwen3.8-27B | — | 219.8 |
| Qwen3-1.7B | 198.1 | — |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B | — | 197.0 |
| MiniCPM5-1B | 125.9 | — |
| Trinity-Mini | 109.0 | — |
| Phi-4-mini-reasoning | 108.2 | — |
| gemma-4-26B-A4B-it-qat | 95.3 | — |
| Qwen3-30B-A3B-Thinking-2507 | 94.3 | — |
| gpt-oss-20b | 93.4 | — |
| OpenReasoning-Nemotron-7B | 68.0 | — |
| WebWorld-8B | 62.3 | — |
| gemma-4-26B-A4B-it | 62.1 | — |
| Carnice-9b | 56.1 | — |
| Muse-Glimmer-30B | 56.0 | — |
| Qwen3.8-Flash-Next-REAP-256-duo | — | 53.5 |
| NVIDIA-Nemotron-Nano-12B-v2 | 40.4 | — |
| Qwen3.8-Flash-Next | — | 40.1 |
| gemma-3-12b-it | 39.4 | — |
| Qwen3-14B | 35.8 | — |
| Nemotron-Cascade-14B-Thinking | 35.7 | — |
| Phi-4-reasoning | 35.4 | — |
| WebWorld-14B | 35.4 | — |
| InternVL3-14B | 35.1 | — |
| DeepSeek-R1-Distill-Qwen-14B | 35.1 | — |
| phi-4 | 34.2 | — |
| DeepHermes-3-Mistral-24B-Preview | 24.5 | — |
| reka-flash-3 | 23.9 | — |
| mistralai_Mistral-Small-3.2-24B-Instruct-2506 | 21.8 | — |
| Tess-4-27B | 18.8 | — |
| GLM-Z1-32B-0414 | 16.6 | — |
| Olmo-3-32B-Think | 16.4 | — |
| gemma-4-31B-it | 15.8 | — |