X local-model bench digest — Sep 18, 2026
As of Sep 18, 2026 — 12 posts collected from X. 30 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 18, 2026 11:23 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 18 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
T4 × Bonsai 2 27B
Sep 18, 2026 · by @yosoymario91
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | llama.cpp | 14.8 tok/s | — |
Tried this with Bonsai 2 27B on a free T4 today, as for some reason v100s are not available for me on Colab. Got PQ2_0 GGUF running with Prism’s llama.cpp build and got ~14.8 tok/s decode, ~3.7s TTFT, and ~9.3 GB VRAM usage. Pretty cool to see a 27B model running this comfortably on a T4. Definitely gonna play around more with this today. Try it, it's completely free.
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, measurement method).
Unspecified platform × GLM-5.3
Sep 18, 2026 · by @LuminaBench · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1M | — | — | 200.0 tok/s | — |
🚨 https://Z.ai just dropped GLM-5.3 FlashX From what I can tell via Vercel it’s basically a really fast version of 5.3 Flash: • up to 200 tok/s • 1M context • native multimodality They’re apparently running this on 100,000 Chinese chips too, so pretty cool https://t.co/cYu4rpBID3
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
M3 Max × Qwen
Sep 18, 2026 · by @AveryChing · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 30.0 tok/s | — |
Sub 10 GB of memory and >30 tok/s on my M3 Max. Still drains the battery like crazy, but nice work @PrismML! Much more usable than Qwen 3.8 27B.
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
M5 Max × Qwen3.8-27B Ternary Bonsai 2
Sep 18, 2026 · by @NFT_Chen
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | — | mlx | 45.0 tok/s | — |
| 262K | — | — | 143.0 tok/s | — |
💥炸了!满血Qwen 3.8 27B 被塞进 Mac 本地就能跑:体积只剩 5.9GB,智商几乎没掉! PrismML 发布 Ternary Bonsai 2 27B:Qwen3.8-27B 血统,三值量化到约 1.76bit,只有满精度的 1/9,综合基准保住 98.2%。不是云端演示,是笔记本直接开跑。 关键突破: 🔹权重仅 5.9GB,不是常见 Q4 的 17–18GB 🔹总分 83.9 vs Qwen3.8 的 85.4(保留 98.2%) 🔹指令跟随 102%,数学 99.5%,代码 99.3% 🔹智能密度 0.444/GB,甩开同级 Q4 / IQ2 / FP16 🔹262K 上下文 + 图文多模态 + Apache 2.0 🔹Apple MLX 可跑,M5 Max 约 45 tok/s;RTX 5090 最高 143 tok/s 以前 27B 要 32GB 统一内存才舒服。现在一台普通…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
Mac Studio M4 Max 128 GB × Bonsai 2 27B
Sep 18, 2026 · by @WescheNex1q
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | mlx | 34.5 tok/s | 242.0 tok/s |
| — | — | mlx | 34.3 tok/s | 242.0 tok/s |
| — | — | mlx | 11.0 tok/s | 33.0 tok/s |
Bonsai 2 27B Mac recipe: Studio & Mini PrismML's ternary quant of Qwen3.8-27B, 1.72 bits/weight. Measured today, same harness as my full-size Qwen3.8-27B recipe: Mac Studio M4 Max 128 GB • 34.5 tok/s prose, 34.3 code, 242 tok/s prefill • Spark Bench 82.4/100, grade B Mac mini M4 24 GB (base, 10-core GPU) • 11.0 tok/s decode, 33 tok/s prefill • Same score trajectory as the Studio The 3.1× decode gap is exactly the…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark × DeepSeek-v4.1-Flash
Sep 18, 2026 · by @WescheNex1q
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | FP8 | vllm | 51.8 tok/s | — |
| — | EXL3 | vllm | 23.7 tok/s | — |
DeepSeek-v4.1 TP4 vs GLM-5.3-Flash Blender Battle: Dodecahedron trapped inside a dodecahedron. One shot, thinking on, temp 0.6, no token cap, no retries. Script runs headless in Blender 5.2. I check the mesh independently. DeepSeek V4.1 Flash: FP8, vLLM TP4, 4 DGX Sparks • 20 min 15 s · 62,928 tokens (59,315 thinking) · 51.8 tok/s • 30 tubes, 20 spheres, 0 intersections • Computed clearance 0.238 — measured 0.238…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state).
M5 Max × ternary 27B
Sep 18, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 55.0 tok/s | — |
A 27B model that fits in under 6GB and runs at 55 tok/s on an M5 Max is a pretty wild local AI package. The new ternary 27B model is based on Qwen3.8-27B and reportedly retains 98.2% of its parent model’s aggregate performance. That’s a huge compression ratio without giving up most of the capability. And the model isn’t being optimized just for simple chat. It’s designed to perform well on workloads like: Agentic…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark x1 × Qwen3.8-Flash-Next
Sep 18, 2026 · by @ViC305
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 240K | EXL3 | exllamav2 523ecd3 | 72.0 tok/s | — |
| 240K | EXL3 | exllamav2 523ecd3 | — | 1150.0 tok/s |
| 4K | EXL3 | exllamav2 523ecd3 | 69.8 tok/s | — |
| 128K | EXL3 | exllamav2 523ecd3 | 69.5 tok/s | — |
| 240K | EXL3 | exllamav2 523ecd3 | 72.0 tok/s | — |
Qwen3.8-Flash-Next - 79.95 tok/s. 🔥 ONE DGX Spark. Full 262K cache configured. 🚀 Qwen3.8-Flash-Next EXL3 just got another major update. New measured default: MTP ndt=5 DSpark, dc=0.6 8-bit KV 262,144-token cache At an actual 240K-token prompt: 72.0 tok/s decode ~1,150 tok/s prefill Exact needle retrieval 𝗙𝗣𝟭𝟲 𝗞𝗩 → 𝟴-𝗕𝗜𝗧 𝗞𝗩 4K context: 66.6 → 69.8 tok/s 128K: 66.6 → 69.5 tok/s 240K: 65.1 → 72.0 tok/s 8-bit KV wins…
Full methodology stated (engine, quantization, context, hardware state, method) — 3 added to the benchmark board (3 of 5 measured rows queued; the rest are published here only).
Unspecified platform × Qwen
Sep 18, 2026 · by @StevensJoe11 · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | llama.cpp 6.4 | 52.0 tok/s | — |
🚨🚨🚨👀 Agentic AI just landed on the edge. NVIDIA ran a 27B Qwen model on a single Jetson AGX Thor at 52 tok/s 6.4× faster than llama.cpp. Multi turn agents, tool use, long context. On-device. That’s the workload Theta EdgeCloud was built for 🔥 Low latency inference across a hybrid EdgeCloud GPU network and not another round trip to a datacenter. Robots, vehicles, factories don’t wait on the cloud... LFG 🚀🌕
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (quantization, context length, hardware state, measurement method).
DGX Spark × Qwen3.8-Flash-Next
Sep 17, 2026 · by @yume_arasaki
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 48.7 tok/s | — |
| — | — | — | 116.0 tok/s | — |
| — | — | — | 60.6 tok/s | — |
| 254K | — | — | 37.0 tok/s | — |
| 122902 | — | — | 43.0 tok/s | — |
People called a DGX Spark an E-Brick. The box is going up in price anyway. Memory bandwidth around 273 GB/s. 4X RTX 6000 Pro guys laughs, saying "Local AI is affordable" at $60,000 setups. So i'm here to show why people are still buying em. Here is everything that runs on ONE Spark right now. 128GB unified. No second node. No cluster. If you bought one and parked it, you are sitting on a studio. INTELLIGENCE: Qwen…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
DGX Spark
Sep 17, 2026 · by @volatilemarkts · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 241K | — | vllm | 1100.0 tok/s | — |
IT WORKS! For a year, anyone with DGX Sparks and Mac Studios has lived with the same problem: the NVIDIA boxes are fast at reading, the Apple boxes are fast at answering, and they can't share a thought. Two islands. A 10 GbE cable between them. Earlier this month we published pd-bridge (https://t.co/GVsHqq9MHC): prefill/decode disaggregation across NVIDIA and Apple silicon. Spark 1 and Spark 2 run the prefill, a Mac…
Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
RTX 4090 × Qwen3.8-27B Unleashed UD-Q3_K_XL
Sep 17, 2026 · by @outsource_
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 250K | Q3_K_XL | dflash2 | 92.0 tok/s | — |
| 128K | Q3_K_XL | dflash2 | 109.0 tok/s | — |
| 64K | Q3_K_XL | dflash2 | 131.0 tok/s | — |
| 8K | Q3_K_XL | dflash2 | 133.0 tok/s | — |
| — | Q3_K_XL | dflash2 | 134.0 tok/s | — |
| 24K | Q3_K_XL | dflash2 | 98.5 tok/s | — |
| 250K | Q3_K_XL | dflash2 | 80.0 tok/s | — |
My First Model On @huggingface Qwen 3.8 27B Unleashed HIT 105,666 Downloads 🔥 Qwen3.8-27B Unleashed UD-Q3_K_XL hits near 100 tk/s @ 250k ctx on 1 x 4090. - DFlash2 - LoopSpec - our async-Q4 Ada kernel - Q4 KV cache - full GPU offload Context Decode Speed ━━━━━━━━━ Short 134 tok/s ───────── 8K 133 tok/s ───────── 64K 131 tok/s ───────── 128K 109 tok/s ───────── 250K 92 tok/s In the actual Hermes harness: - ~24K…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @yosoymario91 — Sep 18, 2026
- @LuminaBench — Sep 18, 2026
- @AveryChing — Sep 18, 2026
- @NFT_Chen — Sep 18, 2026
- @WescheNex1q — Sep 18, 2026
- @WescheNex1q — Sep 18, 2026
- @Oluwaphilemon1 — Sep 18, 2026
- @ViC305 — Sep 18, 2026
- @StevensJoe11 — Sep 18, 2026
- @yume_arasaki — Sep 17, 2026
- @volatilemarkts — Sep 17, 2026
- @outsource_ — Sep 17, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.