X local-model bench digest — Oct 3, 2026
As of Oct 3, 2026 — 12 posts collected from X. 53 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Oct 3, 2026 13:14 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 14 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
RTX PRO 6000 x2 × GLM 5.3 Flash
Oct 3, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | tensorfold | — | 8000.0 tok/s |
| — | EXL3 | tensorfold | 280.0 tok/s | — |
GLM 5.3 Flash EXL3 with TensorFold rocking on 2x RTX PRO 6000 8K tok/s prefill and ~280 tok/s decode 🤯
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 3090 × Qwen3.8-27B
Oct 3, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 25K | — | vllm | 381.0 tok/s | — |
| 25K | — | vllm | 260.0 tok/s | — |
| — | — | vllm | 133.0 tok/s | — |
Qwen3.8-27B just hit around 381 tok/s on a single RTX 3090. 24GB of VRAM. 250W. But there is a very important detail behind that number. This isn’t normal-chat generation. The breakthrough came from changing how speculative decoding uses the information already sitting inside the prompt. The progression is pretty wild: ~82 tok/s Initial single-user setup ~114 tok/s Optimized MTP ~138 tok/s DFlash2 + lookup-augmented…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization).
RTX 170HX x4 × GLM 5.3 Flash
Oct 3, 2026 · by @Morrowmake
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 481.0 tok/s | — |
| — | — | — | 949.0 tok/s | — |
| — | — | — | — | 3197.0 tok/s |
GLM 5.3 Flash on 4 x 170HX - Need for SPEED edition 🔥 Single stream: Up to 481 tok/s (+10%) Concurrency 8: Up to 949 tok/s (+13%) Prefill: 3197 tok/s (+4%) / KV Cache: 1.2 million (+12%) RTX Pro 6000 tier performance at a fraction of the price? Yes please! Full repo below. 👇 https://t.co/rmJXIxT5UA
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x4 × GLM-5.3
Oct 3, 2026 · by @majewskizby
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 32K | — | sparkdash | 33.0 tok/s | — |
| 32K | — | rigmark | 25.4 tok/s | — |
| 32K | — | sparkdash | 44.9 tok/s | — |
| 32K | — | sparkdash | 53.4 tok/s | — |
| 32K | — | sparkdash | 62.4 tok/s | — |
Full GLM-5.3 (753B) TP4 on 4× DGX Spark: recipe update 🦌 • Prose decode: ~33 tok/s at c1 (sparkDash), 25.4 with thinking on (RigMark) • Aggregate prose: c2 44.9, c3 53.4, c4 62.4 tok/s • Context: 32K → 64K • Prefix caching on: repeated prompt TTFT 2.6 s → 0.5 s • Decode drop: ~1% at 30K, ~4% at 60K • Boot: ~470 s → ~284 s What's new: confidence-based early stop for MTP drafts (K1–3 per step), the display memory…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).
M4 Max × Qwen3.8-27B
Oct 3, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | 4-bit | tensorfold | 154.0 tok/s | — |
| — | 4-bit | mlx 0.7.0 | 146.0 tok/s | — |
| — | 4-bit | mlx 26.10.1 | 146.0 tok/s | — |
| — | 4-bit | tensorfold | 114.0 tok/s | — |
| — | 4-bit | mlx 0.7.0 | 111.0 tok/s | — |
| — | 4-bit | mlx 26.10.1 | 115.0 tok/s | — |
| — | 4-bit | tensorfold | 146.0 tok/s | — |
| — | 4-bit | mlx 0.7.0 | 144.0 tok/s | — |
| — | 4-bit | mlx 26.10.1 | 139.0 tok/s | — |
| — | 4-bit | tensorfold | 51.0 tok/s | — |
| — | 4-bit | mlx 0.7.0 | 53.0 tok/s | — |
| — | 4-bit | mlx 26.10.1 | 54.0 tok/s | — |
Three Mac inference engines. One Qwen3.8-27B. One drafter. And the results are basically a tie. Same M4 Max. Same Qwen3.8-27B 4-bit model. Same prompts. Same client. Thinking OFF. The three stacks: TensorFold Numbers: 154 tok/s Code: 114 tok/s JSON: 146 tok/s Story: 51 tok/s oMLX 0.7.0 Numbers: 146 tok/s Code: 111 tok/s JSON: 144 tok/s Story: 53 tok/s mlx-serve 26.10.1 Numbers: 146 tok/s Code: 115 tok/s JSON: 139…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length).
DGX Spark x3 × GLM 5.3
Oct 3, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | tensorfold | 26.0 tok/s | — |
Full GLM 5.3 running on 3x DGX Sparks with TensorFold Currently 24-28 tok/s on prose https://t.co/pKJybzwtL7
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x2 × DeepSeek-V4.1-Flash
Oct 3, 2026 · by @jayleaton
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | tensorfold Tonight's build | 83.0 tok/s | — |
| — | — | tensorfold Tonight's build | 93.5 tok/s | — |
| — | — | tensorfold Tonight's build | 80.0 tok/s | — |
| — | — | tensorfold Tonight's build | — | 1951.0 tok/s |
DeepSeek 4.1 Flash on TensorFold, 2x DGX Spark. Early WIP version. Across-the-board average is 2.2x faster Tonight's build vs vLLM kit from @MiaAI_lab Single stream: 83 vs 32.2 tok/s (2.6x) 4 streams: 93.5 vs 37.6 (2.5x) Code: 79-81 vs 42-45 (1.9x) Prompt Processing: 1833-2069 vs 1070 (1.8-2x) Same answers, just faster: 99.6% agreement with the kit, MMLU 87.5% (the original kit's own score), structured output 22/22…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).
M4 Pro MacBook + iPhone 17 Pro Max × Qwen 3.8 27B
Oct 3, 2026 · by @loudchirper
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 8K | — | llama.cpp mainline | — | 132.0 tok/s |
| 8K | — | llama.cpp mainline | — | 177.0 tok/s |
| 16K | — | llama.cpp mainline | — | 109.0 tok/s |
| 16K | — | llama.cpp mainline | — | 157.0 tok/s |
| 32K | — | llama.cpp mainline | — | 101.0 tok/s |
| 32K | — | llama.cpp mainline | — | 130.0 tok/s |
| 48K | — | llama.cpp mainline | — | 87.0 tok/s |
| 48K | — | llama.cpp mainline | — | 113.0 tok/s |
Upstream runtime flags will not save a 24 GB laptop when a 27B model hits physical memory limits. Mainline llama.cpp PRs can optimize multi-token prediction and tune batch flags all day, but software flags cannot manufacture unified memory out of thin air. When you run Qwen 3.8 27B on a 24 GB M4 Pro MacBook, 64k of 8-bit context fills the machine completely. The workaround here was not waiting for an upstream merge…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
Pixel 10 Pro XL × Gemma 4 E4B
Oct 3, 2026 · by @BlockInsight214
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | litert-lm | 11.0 tok/s | — |
把 Pixel 手机变成一台 OpenAI 服务器 有人把 Pixel 10 Pro XL 直接变成了 AI 服务器:Android 应用,标准 /v1/chat/completions 接口 + 流式输出,TypingMind 等 OpenAI 客户端直连。手机端完全离线跑 Gemma 4 E4B(LiteRT-LM GPU 后端,模型 2.97GB)。实测解码约 11 tok/s,多轮对话复用 KV 前缀,后续轮次约 1.7 秒。三种访问模式:仅本机 / 局域网 / Tailscale 加密。Apache 2.0 开源。
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark x1 × Qwen3.8-Flash-Next
Oct 3, 2026 · by @yume_arasaki
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | custom engine with recoverssm v5.1 | 81.7 tok/s | — |
| — | — | custom engine with recoverssm v5.1 | 146.2 tok/s | — |
| — | — | custom engine with recoverssm v5.1 | 244.4 tok/s | — |
| 8K | — | custom engine with recoverssm v5.1 | 48.6 tok/s | — |
| 32K | — | custom engine with recoverssm v5.1 | 48.4 tok/s | — |
| 131K | — | custom engine with recoverssm v5.1 | 50.2 tok/s | — |
| 262K | — | custom engine with recoverssm v5.1 | 49.4 tok/s | — |
| 35K | — | custom engine with recoverssm v5.1 | 58.2 tok/s | — |
| 35K | — | custom engine with recoverssm v5.1 | 89.6 tok/s | — |
| 35K | — | custom engine with recoverssm v5.1 | 134.3 tok/s | — |
| — | — | custom engine with recoverssm v5.1 | 50.7 tok/s | — |
| — | — | custom engine with recoverssm v5.1 | 122.9 tok/s | — |
The DGX Spark has gone up in price, and that is the bad news. At a new price of $7000 , alot of people are going to ask - can you do much with just 1 DGX Spark? The good news is you can run more frontier-class stuff on it than ever. 60-70 tok/s (322 in concurrency) on Qwen 3.8 Flash is what you can do on it. Yup, you read that right. I spent last night testing the newest proof: @vr8vr8 's single-Spark recipe for…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization).
DGX Spark × Qwen3.8-Flash-Next
Oct 3, 2026 · by @UtaAoya · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 40.0 tok/s | — |
約1年経過したので改めて DGX Spark 、 Mac Book/studio と EVO-X2(Ryzen AI Max+ 395(Strix Halo))の違いを書いておきます。単体でQwen3.8-Flash-Next も40 tok/s で動きます。 1番痛いのはクラスタリングが発展してない所かな〜。2台接続してGLM-5.3-Flashが動けばいいのにね(実験している人はいるみたい)
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Mac Studio M3 Ultra × GLM-5.3-Flash
Oct 3, 2026 · by @volatilemarkts
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256 | 4-bit | tensorfold 0.6.2 | 78.0 tok/s | — |
8 concurrent streams on one M3 Mac Studio. Watch what happens. A language model writes one word-piece per pass through 320 billion numbers. That pass costs about the same for one request or for eight. MLX (oMLX) on my M3 Mac Studio answered them one at a time. Send 8 prompts together and request 8 waits 43 seconds for its first word. All 8 done at 48 s. TensorFold 0.6.2 writes the next piece for all 8 in the same…
Full methodology stated (engine, quantization, context, hardware state, method) — 1 added to the benchmark board (1 of 3 measured rows queued; the rest are published here only).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @MiaAI_lab — Oct 3, 2026
- @Oluwaphilemon1 — Oct 3, 2026
- @Morrowmake — Oct 3, 2026
- @majewskizby — Oct 3, 2026
- @Oluwaphilemon1 — Oct 3, 2026
- @MiaAI_lab — Oct 3, 2026
- @jayleaton — Oct 3, 2026
- @loudchirper — Oct 3, 2026
- @BlockInsight214 — Oct 3, 2026
- @yume_arasaki — Oct 3, 2026
- @UtaAoya — Oct 3, 2026
- @volatilemarkts — Oct 3, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.