DiggTop

LFM2.5-1.2B-Instruct

Llama family · 1.0B params

1recorded run
709.5best tok/s
1verified
Fastest GPUs for this model (best tok/s)
  1. RTX 3090709.5

All benchmarks Methodology

GPUModelQuantBackendCtx tok/sprompt tok/sVRAM GBSource
RTX 3090 · 24 GBLFM2.5-1.2B-InstructQ4_K_Mllama.cpp 0.3.0-dev (build 1, commit d7bd3bf)24576709.5 11527.52.4src