LFM2.5-1.2B-Instruct
Llama family · 1.0B params
1recorded run
709.5best tok/s
1verified
Fastest GPUs for this model (best tok/s)
- RTX 3090709.5
All benchmarks Methodology
| GPU | Model | Quant | Backend | Ctx |
tok/s | prompt tok/s | VRAM GB | Source |
| RTX 3090 · 24 GB | LFM2.5-1.2B-Instruct | Q4_K_M | llama.cpp 0.3.0-dev (build 1, commit d7bd3bf) | 24576 | 709.5 ✓ | 11527.5 | 2.4 | src |