granite-4.2-8b
Llama family · 8.0B params
1recorded run
52.5best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 306052.5
All benchmarks Methodology
| GPU | Model | Quant | Backend | Ctx |
tok/s | prompt tok/s | VRAM GB | Source |
| RTX 3060 · 12 GB | granite-4.2-8b | Q5_K_M | llama.cpp 0.1.2-dev (build 197, commit b062ba7… | 4096 | 52.5 ✓ | 1708.0 | 6.6 | src |