DiggTop

X local-model bench digest — Sep 21, 2026

As of Sep 21, 2026 — 12 posts collected from X. 40 measured runs catalogued, 2 queued for review, 10 kept to the digest only. Latest source post Sep 21, 2026 07:28 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 4 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

RTX 3090 × Qwen3.8-27B

Sep 21, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
hyperqwen 250.0 tok/s
llama.cpp 700.0 tok/s

Qwen3.8-27B spent 21 days trying to build its own CUDA inference engine on a single RTX 3090. Mostly unsupervised. No human-written CUDA. Only around 12 human nudges across the entire experiment. The model wasn’t just given one giant prompt and left alone either. DeepSeek Harness managed the whole operation: Subagents. Roles. Handoffs. Context management. Compaction. HyperQwen served Qwen3.8-27B through vLLM while…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 3090 × Qwen3.8-27B

Sep 21, 2026 · by @mr_r0b0t

Context Quant Engine Decode Prefill
262144 EXL3 4.00bpw buun-llama-cpp 1d0f493c73817176f6953069b7c10211de9555a1 149.2 tok/s
262144 EXL3 4.00bpw buun-llama-cpp 1d0f493c73817176f6953069b7c10211de9555a1 138.7 tok/s
150K EXL3 4.00bpw buun-llama-cpp 1d0f493c73817176f6953069b7c10211de9555a1 27.8 tok/s 525.3 tok/s

@spiritbuun wasn’t kidding about buun-llama serving EXL3! r0b0tlab/Qwen3.8-27B @ 4bpw EXL3 + DFlash2 on a single 24GB RTX 3090: ⚡ 149 tok/s GSM8K 🚀 125 tok/s Q200 🎯 95% Q200 🧠 262K context Runtime: https://github.com/r0b0tlab/buun-llama-qwen3.8-27b-exl3-4.00bpw

Full methodology stated (engine, quantization, context, hardware state, method) — 2 added to the benchmark board (2 of 3 measured rows queued; the rest are published here only).

M3 Pro × K2-Horizon-3.7B

Sep 21, 2026 · by @TeksEdge

Context Quant Engine Decode Prefill
64K Q4_K_M llama.cpp 35.0 tok/s
256K llama.cpp 36.0 tok/s

K2 Horizon got another llama.cpp-family path, this time through TurboQuant. Spark-X2.5 already landed in upstream llama.cpp. I’ve actually got it running myself. A brand-new K2 Horizon TurboQuant PR adds … 🧠 K2 Horizon model support 📦 Hugging Face → GGUF conversion 💬 tokenizer + chat templates ⚙️ llama.cpp-style inference And they’ve already tested 🧠 K2-Horizon-3.7B Q4_K_M 🍎 M3 Pro / 18GB ⚡ ~35 tok/s 📚 64K…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

M4 Max 128GB × Qwen3.8-27B

Sep 21, 2026 · by @Michaelzsguo

Context Quant Engine Decode Prefill
4-bit mlx 73.0 tok/s
mlx 80.0 tok/s

如果你用 Mac 跑本地模型,最近很值得关注的一个项目是 @ddalcu 的mlx-serve。 如果说:Splash 是围绕每个模型重做 inference engine,速度还能再快多少;那么 mlx-serve 走的是另一条路:一套原生 Mac inference engine,能不能把尽可能多的本地模型跑起来? mlx-serve 不是简单地把 MLX 包一层 API, 它的 Server 本身用 Zig 原生实现,不依赖 Python,底层针对 Apple Silicon 做了大量专门优化。它既能直接跑 MLX checkpoint,也把 llama.cpp 集成进来支持 GGUF。不同架构会走不同的 inference path,而不是所有模型都硬塞进同一个 runtime。 它对外同时提供 OpenAI、Anthropic 和 Ollama 兼容接口。所以 Claude…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark x1 × Qwen3.8-Flash

Sep 21, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
117.0 tok/s
180.0 tok/s

Qwen3.8-Flash just got a pretty serious update for running on a single NVIDIA DGX Spark. The new build is hitting: 117 tok/s for prose 180 tok/s for code 8 concurrent streams And that’s on one Spark. The update isn’t just about raw throughput either. There’s now optional official NVIDIA NVFP4 support, giving you another way to optimize the model for the hardware. Peak memory usage has also dropped from around 101…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 3060 × Qwen3.8-27B

Sep 21, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262K llama.cpp PrismML fork 26.0 tok/s
262K llama.cpp PrismML fork 22.0 tok/s
262K llama.cpp PrismML fork 18.0 tok/s
262K llama.cpp PrismML fork 13.0 tok/s

Qwen3.8-27B is building on an RTX 3060 12GB. For two days, the timeline kept saying Bonsai 2 couldn’t actually build anything. So here’s the receipt. 1 hour 24 minutes. One shot. One paragraph of instructions. One RTX 3060 12GB. The full run is shown at 8x speed so you can watch the entire build. This is PrismML’s 5.9GB ternary compression of Qwen3.8-27B, running through the PrismML llama.cpp fork with the full 262K…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

DGX Spark x16 × Kimi K3

Sep 20, 2026 · by @ciprianveg

Context Quant Engine Decode Prefill
vllm 29.8 tok/s
vllm 42.0 tok/s
vllm 58.0 tok/s
vllm 87.1 tok/s
300K vllm 20.0 tok/s

This is the 2.8T-parameter Kimi K3 running at ~30 tok/s on 16× NVIDIA DGX Spark. Not a synthetic “it boots” test. 👇 Actually generating. Usable speed. A few weeks ago, the question was: Can a model this big even run properly on a cluster of Sparks? Then: Can we make it fast enough to actually use? What about for agentic frameworks? I think we're well past that question now. With the new V5, we're…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark x4 × GLM-5.3-Flash-NVFP4

Sep 20, 2026 · by @Tech2Wild

Context Quant Engine Decode Prefill
500K NVFP4+FP8 nvidia/glm-5.3-flash-nvfp4 61.7 tok/s
500K NVFP4+FP8 nvidia/glm-5.3-flash-nvfp4 1997.0 tok/s

🚀 Now running the official NVIDIA quant on 4x DGX Spark (TP4) Rebuilt the NVFP4 attention work on nvidia/GLM-5.3-Flash-NVFP4 instead of LibertAI. Official weights, same speeds. 📊 vs our previous lane: 9 of 9 categories within measurement noise ⚡ Concurrency actually better: +4.8% at C12 → +10.0% at C32 🧠 500K ctx, fp8 KV, 3.53M token pool, cold prefill 1,997 tok/s, acceptance 0.394 → 0.396 ⚔️ vs DeepSeek-V4.1-Flash

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state).

Spark × Qwen3.8-Flash-Next

Sep 20, 2026 · by @MichaelGannotti

Context Quant Engine Decode Prefill
EXL3 3.05bpw exllamav2 79.5 tok/s

Qwen3.8-Flash-Next EXL3 on one Spark: native engine, measured We stood vcruz305's ExLlamaV3 recipe for turboderp's 3.05 bpw Qwen3.8-Flash-Next pack on spark-d369. After a cold-kernel first pass, code decode landed at 79.53 tok/s against their published 79. MiniMax H3 on the other Spark stayed up. https://t.co/ZNjF7wFq0n

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 4090 × Qwen

Sep 20, 2026 · by @outsource_ · third-party report

Context Quant Engine Decode Prefill
250K 143.0 tok/s

Build yourself a local inference dashboard. Measure, tweak, repeat. 🔧 I built BenchHub to track prefill, decode, draft acceptance, and real Hermes tool runs—then used that data to keep improving Qwen Unleashed on my RTX 4090. These calls hit 111–143 tok/s raw decode. We’re also testing retrieval and performance out to 250K context. The goal: sustained 100+ tok/s at long context. Every change has to earn its place…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark × Qwen3.8-Flash-Next

Sep 20, 2026 · by @ViC305

Context Quant Engine Decode Prefill
EXL3 exllamav2 102.6 tok/s
400 EXL3 exllamav2 79.5 tok/s
600 EXL3 exllamav2 71.7 tok/s
EXL3 exllamav2 66.2 tok/s

Qwen3.8-Flash-Next EXL3 crossed 100 tok/s on ONE DGX Spark. 🔥 102.6 tok/s on repetitive code in a community retest of my native ExLlamaV3 recipe. And the part I’m happiest about? My ~79 tok/s code result reproduced at 79.5 on their Spark. 𝗧𝗛𝗘𝗜𝗥 𝗕𝗘𝗡𝗖𝗛𝗠𝗔𝗥𝗞 𝗥𝗘𝗦𝗨𝗟𝗧𝗦 → 102.6 tok/s: repetitive Python code → 79.5 tok/s: my 400-token code workload → 71.7 tok/s: a 600-token prose explanation → 66.2 tok/s: JSON catalog…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state).

DGX Spark × Qwen 3.8 Flash

Sep 20, 2026 · by @yume_arasaki

Context Quant Engine Decode Prefill
EXL3 3.05bpw exllamav2 523ecd3 102.6 tok/s
EXL3 3.05bpw exllamav2 523ecd3 107.1 tok/s
600 EXL3 3.05bpw exllamav2 523ecd3 71.7 tok/s
EXL3 3.05bpw exllamav2 523ecd3 66.2 tok/s
8K EXL3 3.05bpw exllamav2 523ecd3 51.5 tok/s
32K EXL3 3.05bpw exllamav2 523ecd3 55.6 tok/s
64K EXL3 3.05bpw exllamav2 523ecd3 67.4 tok/s
110K EXL3 3.05bpw exllamav2 523ecd3 66.0 tok/s
240K EXL3 3.05bpw exllamav2 523ecd3 55.8 tok/s
35K EXL3 3.05bpw exllamav2 523ecd3 58.0 tok/s 980.0 tok/s
2048 EXL3 3.05bpw exllamav2 523ecd3 86.8 tok/s
EXL3 3.05bpw exllamav2 523ecd3 71.7 tok/s

I have to retract what I posted as the fastest speed of Qwen 3.8 Flash on a single DGX Spark. Cruz shipped a custom ExLlamaV3 fork and a native engine path. I ran the same model on the same Spark: 102.6 max, 80 tok/s on code, 70 tok/s on prose. My previous benchmark and community recipes was 48.6 tok/s. I ran my full grid on it so you don't have to. This might now be the best way to run this model, without breaking…

Full methodology stated (engine, quantization, context, hardware state, method) — 8 added to the benchmark board (8 of 12 measured rows queued; the rest are published here only).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.