CMP 170HX vs RTX 3060 x2
Best verified output tok/s per model (single-stream generation). Hover a value for the quant.
CMP 170HX16 verified runs · best 255.9 tok/s
RTX 3060 x229 verified runs · best 198.1 tok/s
| Model | CMP 170HX | RTX 3060 x2 |
|---|---|---|
| Qwen3.8-27B | 255.9 | — |
| Qwen3-1.7B | — | 198.1 |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B | 165.1 | — |
| MiniCPM5-1B | — | 125.9 |
| Trinity-Mini | — | 109.0 |
| Phi-4-mini-reasoning | — | 108.2 |
| gemma-4-26B-A4B-it-qat | — | 95.3 |
| Qwen3-30B-A3B-Thinking-2507 | — | 94.3 |
| gpt-oss-20b | — | 93.4 |
| Qwen3.6-35B-A3B | 87.2 | — |
| Ornith-1.5-35B-A3B | 87.1 | — |
| OpenReasoning-Nemotron-7B | — | 68.0 |
| WebWorld-8B | — | 62.3 |
| gemma-4-26B-A4B-it | — | 62.1 |
| Ornith-1.5-9B | 61.5 | — |
| Carnice-9b | — | 56.1 |
| Muse-Glimmer-30B | — | 56.0 |
| Qwen3.8-27B-MTP | 46.4 | — |
| Qwen3.8-Flash-Next | 46.1 | — |
| Llama-4-Scout-17B-16E-Instruct | 44.5 | — |
| NVIDIA-Nemotron-Nano-12B-v2 | — | 40.4 |
| gemma-3-12b-it | — | 39.4 |
| Qwen3-14B | — | 35.8 |
| Nemotron-Cascade-14B-Thinking | — | 35.7 |
| Phi-4-reasoning | — | 35.4 |
| WebWorld-14B | — | 35.4 |
| InternVL3-14B | — | 35.1 |
| DeepSeek-R1-Distill-Qwen-14B | — | 35.1 |
| phi-4 | — | 34.2 |
| gemma-4-31B-it | 28.3 | 15.8 |
| DeepHermes-3-Mistral-24B-Preview | — | 24.5 |
| reka-flash-3 | — | 23.9 |
| granite-4.2-30b | 23.7 | — |
| mistralai_Mistral-Small-3.2-24B-Instruct-2506 | — | 21.8 |
| Llama-3.3-70B-Instruct | 20.1 | — |
| Qwen3.6-27B | — | 18.9 |
| Tess-4-27B | — | 18.8 |
| GLM-Z1-32B-0414 | — | 16.6 |
| Olmo-3-32B-Think | — | 16.4 |