Llama-3.3-70B-Instruct
Llama family · 71.0B params
1recorded run
20.1best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- CMP 170HX20.1
All benchmarks Methodology
| GPU | Model | Quant | Backend | Ctx |
tok/s | prompt tok/s | VRAM GB | Source |
| CMP 170HX · 64 GB | Llama-3.3-70B-Instruct | NVFP4 | vllm 0.28.0 | 32768 | 20.1 ✓ | 626.2 | 57.4 | src |