Qwen3.5-4B
Qwen family · 4.0B params
1recorded run
78.6best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 306078.6
All benchmarks Methodology
| GPU | Model | Quant | Backend | Ctx |
tok/s | prompt tok/s | VRAM GB | Source |
| RTX 3060 · 12 GB | Qwen3.5-4B | Q4_K_M | llama.cpp b10443 | 32768 | 78.6 ✓ | 2111.9 | 3.5 | src |