DiggTop

Llama-3.3-70B-Instruct

Llama family · 71.0B params

1recorded run
20.1best tok/s
1verified
Fastest GPUs for this model (best tok/s)
  1. CMP 170HX20.1

All benchmarks Methodology

GPUModelQuantBackendCtx tok/sprompt tok/sVRAM GBSource
CMP 170HX · 64 GBLlama-3.3-70B-InstructNVFP4vllm 0.28.03276820.1 626.257.4src