X local-model bench digest — Sep 19, 2026
As of Sep 19, 2026 — 12 posts collected from X. 49 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 19, 2026 13:36 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 13 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
DGX Spark x4 × DeepSeek-V4.1-Flash
Sep 19, 2026 · by @Tech2Wild
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 500K | EXL3 3.5bpw | — | 91.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 74.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | — | 1378.0 tok/s |
| 500K | EXL3 3.5bpw | — | — | 1401.0 tok/s |
| 500K | EXL3 3.5bpw | — | 324.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 88.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 281.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 75.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 55.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 70.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 39.0 tok/s | — |
| 500K | EXL3 3.5bpw | — | 33.0 tok/s | — |
HUGE SPEED RUN ⚡ Performance improvements, DeepSeek V4.1 Flash on 4x DGX Spark (TP=4, EXL3 3.5bpw, 500K ctx): 💻 Code ~8% faster on 2-4 concurrent streams ~91 tok/s single stream · ~74 tok/s per stream at 2-4 · 324 tok/s aggregate at 6 🚀 🧮 Math ~9% faster on 2-4 concurrent streams ~88 tok/s single stream · 281 tok/s aggregate at 6 🧠 Reasoning ~13% faster single stream, ~11% at 2-4 streams ~75 tok/s single stream 📋…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 3060 × Qwen3.8-27B
Sep 19, 2026 · by @sudoingX
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | — | llama.cpp prismml fork | 26.0 tok/s | — |
| 12K | — | llama.cpp prismml fork | 22.0 tok/s | — |
| 35K | — | llama.cpp prismml fork | 18.0 tok/s | — |
| 77K | — | llama.cpp prismml fork | 13.0 tok/s | — |
the timeline spent two days saying bonsai 2 cannot build. here is 1 hour 24 minutes of it building, one shot, from one paragraph, on an rtx 3060 12gb, sped to 8x so you can watch the whole thing. what you are watching is a 5.9gb ternary compression of qwen 3.8 27b, served through the prismml llama.cpp fork with the full 262k window resident, writing a terminal gpu monitor in one html file. it thinks, it calls a tool…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
DGX Spark × Qwen3.8-27B
Sep 19, 2026 · by @eta1ia
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | vllm | 15.9 tok/s | — |
| — | NVFP4 | sglang | 28.3 tok/s | — |
| — | NVFP4 | vllm | 24.4 tok/s | — |
| — | NVFP4 | sglang | 55.2 tok/s | — |
DGX Spark で動かしている vLLM + Qwen3.8-27B NVFP4 + MTP の構成を SGLang + Qwen3.8-27B NVFP4 + DFlash2 に変更したらからり速度が改善した。 ・通常文章: 15.95 tok/s → 28.30 tok/s ・コーディング: 24.44 tok/s → 55.25 tok/s 個人で使うだけだしこれが良さそう。
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
Tesla V100 PCIe 32GB × Qwen3.8-27B
Sep 19, 2026 · by @UtaAoya
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | ninfer v100 | 29.0 tok/s | — |
| — | NVFP4 | dflash2 | 54.0 tok/s | — |
Local LLM - V100環境-150W縛り 自作PCで、まずは素振りベンチマーク。 V100 x 1 (PG500-216 / GV100GL) なので、単体ではNInferが最速らしい?(9/19 チャッピー調べ) コンテキストを伸ばしてのテストはこれから😊 チャッピー解説 Tesla V100 PCIe 32GB(PG500-216 / GV100GL)単体を、150W縛りで速度比較。 Qwen3.8-27B NVFP4 + ninfer-v100で、通常生成(no-spec)の 約29 tok/s に対し、DFlash2では 最大約54 tok/s を確認。 V100を150Wに制限した状態でも約1.86倍まで高速化できたのがポイント。一方、投機デコードは設定によって出力一致性に差が出るため、そこは継続検証中です。 #地元のV100クラブ
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, measurement method).
iPhone 18 Pro Max × Gemma 4 12B
Sep 19, 2026 · by @bijanbowen
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | Q4_0 | — | 14.0 tok/s | — |
| — | Q4_K_M | — | 57.7 tok/s | — |
| — | Q8_0 | — | 135.5 tok/s | — |
iPhone 18 Pro Max Local LLM: Gemma 4 12B (Q4_0): ~14 tok/s Granite 4.0 H Tiny (Q4_K_M): 57.7 tok/s Qwen 3 0.6B (Q8_0): 135.5 tok/s https://t.co/ff7yVxMKqX
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
Mac Studio M4 Max 128GB × Qwen3.8-27B
Sep 19, 2026 · by @WescheNex1q
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1024 | 4-bit | splash | 68.0 tok/s | — |
| 1024 | 4-bit | mlx | 48.9 tok/s | — |
| 1024 | 4-bit | splash | 53.7 tok/s | — |
| 1024 | 4-bit | mlx | 34.7 tok/s | — |
| 1024 | 4-bit | splash | 33.9 tok/s | — |
| 1024 | 4-bit | mlx | 28.4 tok/s | — |
| 1024 | 4-bit | splash | 36.3 tok/s | — |
| 1024 | 4-bit | mlx | 38.3 tok/s | — |
Splash vs mlx-vlm+MTP, Qwen3.8-27B 4-bit Mac Studio M4 Max 128 GB. Stock splash serve, nothing else on the GPU, 5 coding prompts, 1,024-token cap, medians: Code, no reasoning: Splash 68.0 · MLX 48.9 Code, reasoning on: 53.7 · 34.7 32K+ thinking run: 33.9 · 28.4 Short prose: 36.3 · 38.3 Splash is 20–40% faster on code on this box. Ties on prose. Their 74 tok/s is an M5 Pro number; M4 Max lands 54–68 depending on…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).
RTX 3060 × Ternary Bonsai 2 27B
Sep 19, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 7K | — | llama.cpp | 24.4 tok/s | — |
| 12K | — | llama.cpp | 21.9 tok/s | — |
| 35K | — | llama.cpp | 17.8 tok/s | — |
| 77K | — | llama.cpp | 13.0 tok/s | — |
| 41312 | — | llama.cpp | 15.0 tok/s | — |
| 2K | — | llama.cpp | — | 295.0 tok/s |
| 35K | — | llama.cpp | — | 243.0 tok/s |
| — | — | llama.cpp | 26.1 tok/s | — |
Bonsai 2 27B on an RTX 3060 12GB. Full receipt. This might be one of the more useful local AI tests for anyone sitting on a 12 GB GPU. The model is Ternary Bonsai 2 27B, a compressed version of Qwen3.8-27B. And yes, the entire native 262K context window fits on a 12 GB RTX 3060. Here are the numbers from a live server with thinking enabled: Generation speed by context depth 7K: 24.4 tok/s 12K: 21.9 tok/s 35K: 17.8…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).
DGX Spark × GLM-5.3
Sep 18, 2026 · by @mr_r0b0t · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | — | 58.4 tok/s | — |
Cooked up a GLM-5.3-Flash-EXL3–2.32bpw for the 128GB crowd, perfect for your @NVIDIAAI DGX Spark ⚡️ Testing it now, looks like quantized DFlash2-EXL3 is working well! 58.4 tok/s isn’t to be taken as any sort of eval here but it is a very positive signal 🤩 https://t.co/7JFtjPQuD3
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).
A100 × deepseekv4.1-A100-custom
Sep 18, 2026 · by @shi3z
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | litellm | 120.0 tok/s | — |
| 1M | — | litellm | 50.0 tok/s | — |
| — | — | litellm | 80.0 tok/s | — |
| — | — | litellm | 12979.8 tok/s | — |
Just updated the README for deepseekv4.1-A100-custom! 🚀 Added benchmark specs & performance numbers, along with a full configuration and setup guide (including running on multi-A100 setups). Speed&Multi-Agent: 120tok/s 1M Context Robust 50tok/s Balanced 80tok/s jev_mode: 12,979.8 tok/s Change Claude Code backend from cloud to local via liteLLM Check it out here: 👉 https://t.co/P9RxERX0Ut
Digest-only: confidence 0.40 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).
RTX 4090 × Bonsai 2 27B
Sep 18, 2026 · by @studio_yebisu
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 64.6 tok/s | — |
昨日、Xのゆるトークとかあったから試せなかった Qwen3.8-27Bを3値量子化を施したマルチモーダルLLMであるBonsai 2 27Bを早速弊環境で試してみたんだが。 完全にローカルなのに、64.6 tok/s出るもんだから Asset無し(画像とか渡さず、Net接続もさせず)でなんかそれなりのWebページを25秒ぐらいで作りやがったんですね。 VRAMめっちゃ空いてるし。(12GBあれば十分すぎ) 本家のKimi K3が30tok/sとかだから、用途次第ではめっちゃ活用出来るじゃんって感じなんだよな。なんだよこの謎の3値量子化とかいう新技術。聞いたことねぇよ。 まぁ大分久しぶりの #地元のLLM だったもんだから、進化に驚いた、という話。 環境:7950X3D,RTX4090 24GB,DDR5 64GB
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
M4 Max
Sep 18, 2026 · by @rapidmlx · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | mlx v0.14.3 | 34.0 tok/s | — |
🚀 Rapid-MLX v0.14.3 is live! The star: Bonsai 2 Hadamard 27B. A multimodal 27B model packed into just ~8GB, blazing at 34 tok/s on an M4 Max. Full support across CLI, Server, and Desktop for vision, reasoning, and tool calling. Try it: rapid-mlx serve bonsai2-27b-2bit Huge shoutout to @Mieluoxxx (Morgan Woods) for the original Bonsai 2 loader contribution! 🙌 On top of that ⚡️ Multimodal conversations are now much…
Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (quantization, context length, hardware state, measurement method).
DGX Spark × glm-5.3
Sep 18, 2026 · by @0xSero · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | EXL3 | — | 20.0 tok/s | — |
DGX Spark owners rejoice! Finally managed to make a better quant for glm-5.3-flash, you get 262k context, vision, and 10-20 tok/s on 1x DGX Spark. 80% top 1 agreement, lower KL, lower perplexity. Enjoy! Great as an overnight agent https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Spark https://t.co/HmuRMzVroM
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, hardware state, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @Tech2Wild — Sep 19, 2026
- @sudoingX — Sep 19, 2026
- @eta1ia — Sep 19, 2026
- @UtaAoya — Sep 19, 2026
- @bijanbowen — Sep 19, 2026
- @WescheNex1q — Sep 19, 2026
- @Oluwaphilemon1 — Sep 19, 2026
- @mr_r0b0t — Sep 18, 2026
- @shi3z — Sep 18, 2026
- @studio_yebisu — Sep 18, 2026
- @rapidmlx — Sep 18, 2026
- @0xSero — Sep 18, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.