mistralai_Mistral-Small-3.2-24B-Instruct-2506
1recorded run
21.8best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 3060 x221.8
| GPU | Model | Quant | Backend | Ctx | tok/s | prompt tok/s | VRAM GB | Source |
|---|---|---|---|---|---|---|---|---|
| RTX 3060 x2 · 12 GB | mistralai_Mistral-Small-3.2-24B-Instruct-2506 | Q4_K_M | llama.cpp b10443 | 65536 | 21.8 ✓ | 586.5 | 17.4 | src |