Phi-4-mini-reasoning
1recorded run
108.2best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 3060 x2108.2
| GPU | Model | Quant | Backend | Ctx | tok/s | prompt tok/s | VRAM GB | Source |
|---|---|---|---|---|---|---|---|---|
| RTX 3060 x2 · 12 GB | Phi-4-mini-reasoning | Q4_K_M | llama.cpp 13f2b28 | 512 | 108.2 ✓ | 3875.2 | 8.6 | src |