Llama-4-Scout-17B-16E-Instruct
Llama family · 109.0B params
1recorded run
44.5best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- CMP 170HX44.5
All benchmarks Methodology
| GPU | Model | Quant | Backend | Ctx |
tok/s | prompt tok/s | VRAM GB | Source |
| CMP 170HX · 64 GB | Llama-4-Scout-17B-16E-Instruct | UD-Q3_K_XL | llama.cpp 0.28.0 | 2048 | 44.5 ✓ | 584.4 | 46.6 | src |