X local-model bench digest — Sep 23, 2026
As of Sep 23, 2026 — 12 posts collected from X. 23 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 23, 2026 10:18 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 8 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
Unspecified platform × llama.cpp
Sep 23, 2026 · by @onusoz · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 16K | — | llama.cpp | 10.0 tok/s | — |
That old macbook with 8gb unified memory or windows gaming laptop of yours can now be your fully local and private AI server Customer data. Client emails. Health records that you cannot just paste into chatgpt Now you can give them to your local AI however you like! Connect the laptop to your tailscale network Download Bonsai 2 27b and its llama.cpp build (might take some days/weeks for official llama.cpp to support…
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
DGX Spark x1 × DiffusionGemma
Sep 23, 2026 · by @mmastrac
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | vllm | 97.0 tok/s | — |
| — | — | vllm | 317.0 tok/s | — |
DiffusionGemma last experiment before I sleep hit 97 tok/s at conc=1 and 317 tok/s at conc=32. Honestly nuts, that's one spark only. Why was this ignored for so long? vLLM patches landing over next few days.
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
RTX Pro 6000 × MiMo v2.6 Flash
Sep 23, 2026 · by @SSHCodes
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 512K | — | — | 40.0 tok/s | — |
| 512K | — | — | — | 2000.0 tok/s |
I shoved MiMo v2.6 Flash UNQUANTIZED onto ONE RTX pro 6000 and 128gb RAM and after a few days of nonstop work got 40 tok/s decode speed and 2000 tok/s prefill with 512k context window (that is perfectly usable to me!!) Fully open source and reproducible, give it a try! https://t.co/3FQrRmLjRE
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 3080 Ti × Qwen3.8-27B
Sep 23, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 131K | — | — | 47.0 tok/s | — |
“98.2% of Qwen3.8-27B performance retained.” So I gave the model an actual job. My standard Three.js FPS prompt, the same kind of task I normally use to see how a model handles agentic coding. Ran it for 6 hours on an RTX 3090. The first 32K tokens went into planning. No files. Eventually it produced: A black screen. Two shaders that didn’t compile. A player that spawned dead. Then the final report said it was…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 5090 × Ternary Bonsai 2 27B
Sep 23, 2026 · by @xueyu1125
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 143.0 tok/s | — |
| — | — | — | 46.8 tok/s | — |
今天想和大家分享一个8G显存跑视频工作流的最后一块版图,完全基于 Qwen3.8-27B 模型架构的开源权重 Ternary Bonsai 2 27B Bonsai 2 27B 最大的亮点,是把约 54GB 的 Qwen3.8-27B 压到 5.95GB(缩小了9倍) 同时官方测试仍保留 98.2% 的综合能力,尤其数学、编程和工具调用损失很小 在英伟达RTX 5090 可以达到 143 tok/s,在Mac M5 Max上可以达到 46.8 tok/s 这有可能是最适合16G显存 甚至8G显存的本地模型 Bonsai 2 27B 写剧本拆分镜 - Qwen-Image-2.1 生成关键帧 - MiniMax H3 生成视频
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark × MiMo-V2.6-Flash-RL
Sep 23, 2026 · by @benthecarman
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 250K | EXL3 | — | 30.0 tok/s | — |
Got MiMo-V2.6-Flash-RL running on a single DGX Spark. Quantized using EXL3 at about 30 tok/s with a 250K context window, measured only a ~1.5% quality loss. https://huggingface.co/benthecarman/MiMo-V2.6-Flash-RL-exl3
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
DGX Station GB300 × MiMo-V2.6-Pro
Sep 22, 2026 · by @JamesMeadlock
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 11500 | — | — | 34.3 tok/s | — |
| 11500 | — | — | 33.0 tok/s | — |
New recipe: MiMo-V2.6-Pro (1.02T MoE) on one DGX Station GB300. Experts live in Grace RAM. The busiest ones move to HBM, ranked by real agent traffic. Decode 31.2 → 37.4 tok/s at 11.5K context. Tool calls 29.5 → 36.6. Weights as shipped. https://t.co/v6O6ATohnL
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
M3 Ultra × Xiaomi MiMo-V2.6 Flash
Sep 22, 2026 · by @rapidmlx
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | mlx 0.15.0 | 54.0 tok/s | — |
Rapid-MLX 0.15.0 is live! 🚀 Run Xiaomi MiMo-V2.6 Flash locally on high-memory Apple Silicon — a 309B-total / 15B-active MoE reaching 49–59 tok/s on M3 Ultra. Features verified tool calling, structured JSON output, and prompt caching that cuts a 15K context repeat from 34.2s to 0.95s. What’s new in 0.15.0: • Qwen-Image 2.1: Integrated across Server & Desktop (text-to-image & editing) with Mac-optimized memory…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).
DGX Spark x2 × Qwen3.8-Flash
Sep 22, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 146.0 tok/s | — |
| — | — | — | 56.8 tok/s | — |
| — | — | — | — | 3500.0 tok/s |
Great numbers on the new M5 Ultra. Here's Qwen3.8-Flash on 2x DGX Sparks: Prose on 4 streams: 146.0 tok/s Single stream prose - 56.8 tok/s Prefill is 3500 tok/s So basically the same?
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
RTX 3060
Sep 22, 2026 · by @sudoingX
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 32768 | — | prismml | — | 267.0 tok/s |
| 32768 | — | prismml | — | 522.0 tok/s |
update on the same bundle. a new prefill kernel landed in the prismml fork this morning, i rebased onto it and rebuilt, and prompt processing went from 267 to 522 tok/s on the 3060. this one is purely how fast it reads your prompt. a 32k prompt that took two minutes to load now takes one. same weights, same gguf, nothing to re-download except the 492 mb bundle. i ran both builds back to back in one session rather…
Digest-only: model.name missing; incomplete methodology (engine+version, quantization, hardware state).
DGX Spark x2 × GLM-5.3-Flash-NVFP4
Sep 22, 2026 · by @Bizuayeu
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 BIZ AXL | vllm 1.11.3 | 31.4 tok/s | — |
| — | NVFP4 BIZ AXL | vllm 1.11.3 | 41.3 tok/s | — |
| 38962 | NVFP4 BIZ AXL | vllm 1.11.3 | — | 1294.8 tok/s |
GLM-5.3-Flash の 2x-DGX-Sparks 用セットアップ "NVFP4 BIZ AXL" を、遂にまとめ上げることができました。使い易いライセンスにこだわり、標準 MTP の利用を継続したまま、散文 31.38 tok/s(コード 41.28 tok/s)を出せ、256k セッションの並走が可能です。商用制限無し。 https://github.com/Bizuayeu/GLM-5.3-Flash-NVFP4-2x-DGX-Sparks-BIZ/
Full methodology stated.
NVIDIA CMP 170HX x2 × Qwen3.8-Flash-Next-W4A16-FP8PLE
Sep 22, 2026 · by @MoonlitMaven
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262052 | W4A16-FP8PLE | vllm 0.29.0 | 186.9 tok/s | — |
| 262052 | W4A16-FP8PLE | vllm 0.29.0 | — | 4770.6 tok/s |
| 262052 | W4A16-FP8PLE | vllm 0.29.0 | — | 121035.8 tok/s |
PCIe x1 is supposed to be terrible for LLM inference… right? I benchmarked Qwen3.8-Flash-Next-W4A16-FP8PLE on a host with: 2× unlocked NVIDIA CMP 170HX used for inference 65,536 MiB VRAM per GPU 38.04 GiB system RAM Intel Core i5-8500T PCIe Gen2 x1 on both inference GPUs Using vLLM 0.29.0, Pipeline Parallel = 2, MTP = 3: ⚡ 186.93 output tok/s @ 8 concurrent requests 🚀 4,770.57 input tok/s cold prefill at 262,052…
Full methodology stated (engine, quantization, context, hardware state, method) — 1 added to the benchmark board (1 of 3 measured rows queued; the rest are published here only).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @onusoz — Sep 23, 2026
- @mmastrac — Sep 23, 2026
- @SSHCodes — Sep 23, 2026
- @Oluwaphilemon1 — Sep 23, 2026
- @xueyu1125 — Sep 23, 2026
- @benthecarman — Sep 23, 2026
- @JamesMeadlock — Sep 22, 2026
- @rapidmlx — Sep 22, 2026
- @MiaAI_lab — Sep 22, 2026
- @sudoingX — Sep 22, 2026
- @Bizuayeu — Sep 22, 2026
- @MoonlitMaven — Sep 22, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.