Llama-3.2-1B-Instruct
1recorded run
699.9best tok/s
1verified
Fastest GPUs for this model (best tok/s)
| GPU | Model | Quant | Backend | Ctx | tok/s | prompt tok/s | VRAM GB | Source |
|---|---|---|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell · 96 GB | Llama-3.2-1B-Instruct | Q8_0 | llama.cpp 876a432 | 4096 | 699.9 ✓ | 64851.3 | 2.1 | src |