Radeon AI Pro R9700 x3 vs RTX 3060 x2
Best verified output tok/s per model (single-stream generation). Hover a value for the quant.
Radeon AI Pro R9700 x37 verified runs · best 58.8 tok/s
RTX 3060 x229 verified runs · best 198.1 tok/s
| Model | Radeon AI Pro R9700 x3 | RTX 3060 x2 |
|---|---|---|
| Qwen3-1.7B | — | 198.1 |
| MiniCPM5-1B | — | 125.9 |
| Trinity-Mini | — | 109.0 |
| Phi-4-mini-reasoning | — | 108.2 |
| gemma-4-26B-A4B-it-qat | — | 95.3 |
| Qwen3-30B-A3B-Thinking-2507 | — | 94.3 |
| gpt-oss-20b | — | 93.4 |
| OpenReasoning-Nemotron-7B | — | 68.0 |
| WebWorld-8B | — | 62.3 |
| gemma-4-26B-A4B-it | — | 62.1 |
| Qwen3-VL-8B-Instruct | 58.8 | — |
| Carnice-9b | — | 56.1 |
| Muse-Glimmer-30B | 47.8 | 56.0 |
| NVIDIA-Nemotron-Nano-12B-v2 | — | 40.4 |
| gemma-3-12b-it | — | 39.4 |
| Ling-3.0-flash | 39.3 | — |
| Qwen3-14B | — | 35.8 |
| Nemotron-Cascade-14B-Thinking | — | 35.7 |
| Ornith-1.5-35B-A3B | 35.4 | — |
| Phi-4-reasoning | — | 35.4 |
| WebWorld-14B | — | 35.4 |
| InternVL3-14B | — | 35.1 |
| DeepSeek-R1-Distill-Qwen-14B | — | 35.1 |
| phi-4 | — | 34.2 |
| Ornith-1.0-9B | 31.2 | — |
| Qwen3.8-Flash-Next | 28.1 | — |
| DeepHermes-3-Mistral-24B-Preview | — | 24.5 |
| reka-flash-3 | — | 23.9 |
| gemma-4-31B-it | 23.9 | 15.8 |
| mistralai_Mistral-Small-3.2-24B-Instruct-2506 | — | 21.8 |
| Qwen3.6-27B | — | 18.9 |
| Tess-4-27B | — | 18.8 |
| GLM-Z1-32B-0414 | — | 16.6 |
| Olmo-3-32B-Think | — | 16.4 |