Qwen3.8-Flash-Next-REAP-256-duo
1recorded run
53.5best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 3090 x253.5
| GPU | Model | Quant | Backend | Ctx | tok/s | prompt tok/s | VRAM GB | Source |
|---|---|---|---|---|---|---|---|---|
| RTX 3090 x2 · 24 GB | Qwen3.8-Flash-Next-REAP-256-duo | Unsloth-Dynamic-Q3_K_XL | llama.cpp 0.3.0-dev / commit 035e227 (llamacpp… | 32768 | 53.5 ✓ | 948.6 | 33.9 | src |