DiggTop

X local-model bench digest — Sep 20, 2026

As of Sep 20, 2026 — 12 posts collected from X. 36 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 20, 2026 12:27 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 9 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x3 × GLM 5.3 Flash

Sep 20, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
EXL3 95.0 tok/s

95 tok/s 🤯 GLM 5.3 Flash EXL3 running 4 concurrent streams on 3x DGX Sparks. And no, this is NOT "structured" decode. This is mostly prose with tiny bit of code in mid chat. A dream! https://t.co/EPTIsFXKFX

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark x1 × Qwen3.8-Flash-Next-NVFP4

Sep 20, 2026 · by @Tech2Wild

Context Quant Engine Decode Prefill
262144 NVFP4 vllm 8a728663 43.9 tok/s
262144 NVFP4 vllm 8a728663 32.5 tok/s
262144 NVFP4 vllm 8a728663 21.7 tok/s
262144 NVFP4 vllm 8a728663 44.3 tok/s
262144 NVFP4 vllm 8a728663 68.8 tok/s
7K NVFP4 vllm 8a728663 1269.0 tok/s
28K NVFP4 vllm 8a728663 1757.0 tok/s
113K NVFP4 vllm 8a728663 1760.0 tok/s

If You Got 1 x DGX Spark you should NOT be running Qwen 3.8 27B. You should be running Qwen 3.8 Flash Next. 1M KV POOL and PEAK 43 Tok/s https://github.com/tonyd2wild/Qwen3.8-Flash-Next-NVFP4-DGX-Spark

Full methodology stated (engine, quantization, context, hardware state, method) — 1 added to the benchmark board (1 of 8 measured rows queued; the rest are published here only).

M5 Max MacBook Pro × Qwen3.8-27B

Sep 20, 2026 · by @Michaelzsguo

Context Quant Engine Decode Prefill
inco splash 50.7 tok/s
inco splash 130.3 tok/s
inco splash 147.1 tok/s
inco splash 165.7 tok/s

这是我目前见过,在 Mac 上跑 Qwen3.8-27B 最快的方案:Inco Splash。 官方在 M5 Max MacBook Pro 上测到最高 144 tok/s 的 Decode。我自己也跑了一遍 benchmark,基本复现了这个量级: Creative fiction:50.7 tok/s Coding:130.3 tok/s Boilerplate code:147.1 tok/s 近乎确定性的重复输出:165.7 tok/s 不同任务的速度差别很大,但在 Coding 和 Agent 常见的输出模式下,130–160 tok/s 确实可以跑出来。 Splash 不是一套追求支持所有模型的通用引擎。它围绕具体的 Qwen 模型和 Apple Silicon 做了深度优化,包括定制的 Metal kernels、8-bit KV cache,以及专门训练的 speculative decoding draft…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash-NVFP4

Sep 20, 2026 · by @Bizuayeu

Context Quant Engine Decode Prefill
256K NVFP4 + W4A16 (attention project vllm 45.4 tok/s
256K NVFP4 + W4A16 (attention project vllm 34.8 tok/s
256K NVFP4 + W4A16 (attention project vllm 24.3 tok/s

そんな車輪の再開発がこちらですw https://github.com/Bizuayeu/GLM-5.3-Flash-NVFP4-2x-DGX-Sparks-BIZ vLLM のバグまで潰して推論ノイズを補正し、ちょっとした追加量子化を行うことで、GLM-5.3-Flash の 2x Sparks での生成速度が最大 45.4 tok/s に届きました。MTP の有効予測深度を k=5 まで上げるのが当面の目標。

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).

RTX 3060 × qwen

Sep 20, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
262K 40.0 tok/s

you can now run qwen 3.8 27b, bonsai 2 at 40 tok/s on a single rtx 3060 or any 12gb card, full agentic loop, the whole 262k context window resident. let that sink in. i am saying this because i ran it myself and built things with it. a 77 minute agent build on that card, a working file at the end, hermes agent on top with 25 tools loaded. i know it sounds impossible, the closed labs have spent two years telling you…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 3090 × Qwen3-30B-A3B

Sep 20, 2026 · by @0x0SojalSec · third-party report

Context Quant Engine Decode Prefill
150.0 tok/s
80.0 tok/s
100.0 tok/s
200.0 tok/s
400.0 tok/s
350.0 tok/s

Best Local AI hardware cheat sheet, Hardware & Budgets : 2026 - For a daily driver at $1.6K: Hardware: Arc B70, RTX 3090 Model: Qwen3-30B-A3B/Qwen3.8-27B/Ornith familly Speed: 50-150 tok/s - For a high-memory setup for Coding, photo & video editing at $3.7K: Hardware: Framework Desktop/Strix Halo Model: Qwen3.8-Next-Flash/ DeepSeek-V4 Flash Speed: 40-80 tok/s - For a CUDA workstation Coding, learning, local stack at…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 3090 × Qwen3.8-27B

Sep 19, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
dflash2 + lookup
dflash2 + lookup 138.0 tok/s
mtp optimized 114.0 tok/s
133.0 tok/s

Qwen3.8-27B just hit 381 tok/s on a single RTX 3090. But there’s an important detail behind that number. This isn’t normal generation speed. The setup is using context lookup, where the model can verify multiple tokens at once when the answer is already present in the prompt. With DFlash2 + lookup, it reaches about 138 tok/s. Optimized MTP gets around 114 tok/s. Then the longer verification setup kicks in. With…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × Qwen 3.8 27B

Sep 19, 2026 · by @antonioleivag

Context Quant Engine Decode Prefill
42.0 tok/s

¡Madre mía, Qwen 3.8 27B en local a 42 tok/s! 🤯 Para que te hagas una idea, hasta ahora solo había conseguido unos 22 tok/s con oMLX o similares, y en LM Studio el modelo estándar con el mismo ejemplo ha ido a 11 tok/s. ¿Como lo he conseguido? Te cuento 👇👇 https://t.co/hR5cLPlsmJ

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Arc Pro B70 × Qwen3.8-27B

Sep 19, 2026 · by @TeksEdge · third-party report

Context Quant Engine Decode Prefill
Q4_K_XL llama.cpp 23.4 tok/s
Q8_K_XL llama.cpp 15.9 tok/s

Intel's 32GB Arc Pro B70 finally got a proper Qwen3.8-27B Local AI workout. So this is why the 32GB VRAM is important, even from an Intel GPU. Here is Gigazine's setup. Windows 11 using Unsloth Desktop + llama.cpp ... 🧠 Qwen3.8-27B Q4_K_XL 💾 30.0GB VRAM + ~1GB shared RAM 🚀 23.4 tps Then they tried Q8: 🧠 Qwen3.8-27B Q8_K_XL 💾 30.2GB VRAM + ~1.5GB shared RAM ⚡ 15.9 tps That's a full 27B-class model running at very…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform × llama.cpp

Sep 19, 2026 · by @SergiiioBS · third-party report

Context Quant Engine Decode Prefill
262K INT4 vllm 96.0 tok/s
262K INT4 vllm 39.0 tok/s

I ran the same 35B MoE model (INT4) on one Arc Pro B70 with 3 different inference engines, from 512 to 262K tokens of context. What I found: - vLLM is fastest at short and medium context (~96 tok/s) but dies past 131K on 32GB - OpenVINO is steady (~30-39 tok/s) up to 64K, then its 131K prefill crashes. Bug filed: openvino.genai#4505 - llama.cpp is the only engine that runs the full ladder to 262K. With the draft…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, hardware state, measurement method).

Sep 19, 2026 · by @toto2AI

Context Quant Engine Decode Prefill
Q8 lmstudio 63.7 tok/s
Q8 lmstudio 82.0 tok/s
lmstudio 50.0 tok/s

『妖精×月見』🌕#AI妖精部 下記はオタクな日記ですので、興味のない方はイラストをお楽しみくださいませ🍡 ローカルLLM日記:外部eGPUを2枚(OCuLink+USB4)つないで 63.66 → 82.01 tok/s(約1.29倍)。USB4は理論値に届かないけど、安くVRAM32GB使えるの嬉しい。NVIDIAドライバでTensor Parallelism使用😊 ✨検索してください。「Nvidiaドライバ LMstudio "Tensor Parallelism"」増速理論値はNVIDIA発表では1.8倍です。 ・USB4という超遅い接続でも遅くなるペナルティが無かったという事が重要です。 ・データについて。63.66トークン毎秒という数値はGPU1枚の時の値。USB4接続のeGPUをマルチで繋いでLMstudioでTensor Parallelism使用で 82.01でした。…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform × Deepseek 4.1 Flash

Sep 19, 2026 · by @DamianCatanzaro · third-party report

Context Quant Engine Decode Prefill
30.0 tok/s

Estoy haciendo unas pruebas con Deepseek 4.1 Flash y necesito que todas las AIs tengan esta cantidad de tok/s, todavía no se si me va a dar el resultado que necesito, pero verlo trabajando a esta velocidad es asombroso comparado con mi día a día a 20/30tok/s máximo. https://t.co/MXw3esWAxT

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.