DiggTop ⌕

X local-model bench digest — Oct 9, 2026

As of Oct 9, 2026 — 32 posts collected from X. 79 measured runs catalogued, 32 kept to the digest only. Latest source post Oct 9, 2026 10:54 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.

Unspecified platform × mistral 4

Oct 9, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
181K — — 94.3 tok/s —

i turned mistral 4 le chonk's thinking off and it built the floating tree still behind glm 5.3 flash and qwen 3.8 flash next on the base and the island, those two built prettier worlds local on my desk, but le chonk writes files now lol 12/12 tests green, visual gate passed 139 calls, 101k output tokens 94.3 tok/s median decode on mistral's cloud $3.10 for the whole build, about 70 minutes 6 problems caught in its…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

Sapphire R9700 × Qwen3.8-27B

Oct 9, 2026 · by @BlockInsight214

Context Quant Engine Decode Prefill
131K Q4 lemonseed 31.0 tok/s —
131K Q4 lemonseed 109.0 tok/s —
131K Q4 lemonseed 159.0 tok/s —
4K Q4 lemonseed — 1636.0 tok/s
32K Q4 lemonseed — 1395.0 tok/s

⚡Qwen3.8-27B 在 iPad 外接显卡上跑出了 159 token/秒 今天刚出的实测:iPad Pro 通过雷电外接 Sapphire R9700 显卡,跑 Qwen3.8-27B(Q4 量化,131K 上下文): 基线解码:31 tok/s 加 MTP 投机解码:109 加 DFlash2 草稿树:159 Strix Halo 核显也能到 64。 背后的 LemonSeed 引擎把模型的前向过程录成图、在当前设备上逐个实测选最快的核,投机解码的草稿数量也按实测接受率动态定。 预填充 4K token 约 1636 tok/s,拉到 32K 还剩 1395。 引擎已开源:https://t.co/AT2nEbLKzo 启示比硬件本身大:同一个模型,解码策略选对,速度白捡 5 倍。

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Unspecified platform × Qwen

Oct 9, 2026 · by @manateelazycat · third-party report

Context Quant Engine Decode Prefill
— — vllm 60.0 tok/s —

国庆在家7天,重新设计了懒猫 AI 算力舱软件平台 全新设计的平台,把底层的算力舱硬件,中间的 AI 模型和上层的 AI 应用完全解耦了 1. 用户可以精确的控制每台算力舱运行的 AI 模型,包括显存、磁盘控制,还可以实时监控每个算力舱的产出 Token,看看到底给用户节省了多少费用 2. 官方优化了主流的 25 个 AI 模型,覆盖 GLM 5.3 Flash、DeepSeek、Qwen 3.8 27B、 Qwen 3.8 Flash Next、MiniMax H3 等,基于 vLLM、SGLang、TensorFold 三套主流框架深度优化, GLM 5.3 Flash 的解码速度高达 60 Tokens / s, 简直就是离线的 AI 生产力神器,最关键的是,所有模型都支持一键自动部署,节省你大量折腾 AI 模型调优的时间,更多的时间用于创作上 3. 开发者可以基于这个算力平台,自由的组合 AI 模型开发你的 AI…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark × Qwen3.8-27B

Oct 9, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— 4-bit tensorfold 0.6.5 69.0 tok/s —
— 4-bit tensorfold 0.6.5 50.0 tok/s —
— 8-bit tensorfold 0.6.5 38.0 tok/s —

Qwen3.8-27B Mac m4 max vs DGX spark: same engine (TensorFold 0.6.5 + DFlash2), same seeds: DGX Spark • 4-bit: 82.9 · 52 min (~69 tok/s) Mac Studio M4 Max • 4-bit: 83.6 · 71 min (~50 tok/s) • 8-bit: 84.7 · 95 min (~38 tok/s) Hardware moved it

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state).

RTX 3090 × Qwen3.8-27B

Oct 9, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
25K — vllm — —
25K — vllm — —
— — vllm 133.0 tok/s —

Qwen3.8-27B has gone from ~82 tok/s to 381 tok/s on a single RTX 3090. Same 24GB GPU. Around 250W. The big change is how the inference stack uses information already sitting in the prompt. Here’s the progression: • ~82 tok/s: initial single-user setup • ~114 tok/s: optimized MTP • ~138 tok/s: DFlash2 + lookup drafting • ~381 tok/s: longer verification blocks + context lookup The setup combines optimized vLLM…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).

RTX 3090 × Qwen3.8-27B

Oct 9, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 — llama.cpp custom fork 69.9 tok/s —
262144 — llama.cpp custom fork — 1124.0 tok/s
60K — llama.cpp custom fork 27.9 tok/s —

Qwen3.8-27B running at nearly 70 tok/s with a 262K context window on a 12GB GPU sounds impossible until you look at the engineering behind it. The setup uses Mirai S Qwen3.8-27B, a heavily compressed 2.4-bit model weighing roughly 11GB, converted to GGUF for a custom llama.cpp fork without re-quantizing the original weights. The initial port ran on an RTX 3090 limited to 12GB VRAM and managed about 38 tok/s with an…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (measurement method).

M5 Ultra × Qwen

Oct 9, 2026 · by @ashen_one · third-party report

Context Quant Engine Decode Prefill
— — — 70.0 tok/s —

After three days of trying, I'm happy to take my L on this I tried my best to run Qwen 3.8 Flash Next 6bit on my 256 GB M5 Ultra for the past 3 days, to see how good it can design with the Blender MCP and integrate it into the Unity CLI I had my Dot running on Astra check in every 10 minutes to ensure that it was working properly and give it reference images, but this is the best I could do I knew the local models…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

M5 Ultra 80-core 256 GB × Qwen 3.8 Flash

Oct 9, 2026 · by @yume_arasaki

Context Quant Engine Decode Prefill
195K — mlx 0.6.4 96.6 tok/s —
— — mlx 0.6.4 83.0 tok/s —
262K — mlx 73.5 tok/s —
— — tensorfold 0.3.4 200.7 tok/s 471.0 tok/s
— — mlx 147.0 tok/s 3450.0 tok/s
— — mlx 160.0 tok/s 3400.0 tok/s
256K — mlx 74.7 tok/s —
— — mlx 108.0 tok/s —
— — tensorfold 0.3.4 200.7 tok/s 471.0 tok/s
— — mlx 226.0 tok/s —
— — mlx 134.0 tok/s —

The M5 Ultra Mac Studios are finally landing on desks, and the Qwen 3.8 Flash numbers are pouring in. Every chip from M4 Pro to M5 Ultra, every number with its source. Here is what I found. I wanted to find out what these machines actually do with one of the world's most popular models, and whether the older chips are still worth running. Note: I stand corrected about the M3 Ultra which I took from Apple's trade-in…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, hardware state, measurement method).

DGX Spark

Oct 8, 2026 · by @Tech2Wild · third-party report

Context Quant Engine Decode Prefill
— — — 174.0 tok/s —

GLM 5.3 Flash Is the Best Local Model on DGX Spark (174 tok/s) https://youtu.be/Br1NklpIkuo https://t.co/c0pAdasfsN

Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark

Oct 8, 2026 · by @landontgreen · third-party report

Context Quant Engine Decode Prefill
— — — 100.0 tok/s —

Almost 100 tok/s PROSE on GLM 5.3 Flash 4x DGX Sparks

Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Station x2 × DeepSeek-V4.1-Flash

Oct 8, 2026 · by @SJRonanMD

Context Quant Engine Decode Prefill
8K — b12x public b12x code 5774.0 tok/s —
8K — b12x public b12x code 4548.0 tok/s —
— — b12x public b12x code 450.0 tok/s —

Deepseek V4.1 Flash at 5774 tok/s Each DGX Station pairs a GB300 with an RTX PRO 6000, and we made the two cards work as one. DeepSeek-V4.1-Flash picks from 384 experts per layer. The 285 most-used experts live in the GB300's HBM. The other 99 live on the RTX PRO 6000, which computes them itself as an expert sidecar. Only the tokens routed to those experts cross PCIe, and if the 6000 ever stalls, the GB300 carries…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization).

M3 Ultra x4 × GLM-5.3

Oct 8, 2026 · by @volatilemarkts

Context Quant Engine Decode Prefill
622K — mlx 8.5 tok/s —
622K — tensorfold new Zig engine 48.0 tok/s —

BACK TO THE LAB AGAIN: @ashxhart We ran his code entirely on a cluster of 4x M3 Ultra Mac Studios serving GLM-5.3. 622k context. https://t.co/VY3KTMKgD5 What went down: TensorFold & GLM-5.3 (753B) on 4x Mac M3 Ultra's. TensorFold’s new Zig engine. One massive model sharded across 4 Macs over a custom Thunderbolt-5 RDMA mesh. 9 TOK/S mlx became 48 tok/s with Tensorfold. A literal 5x improvement. 8-9tok/s MLX improved…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Also in this window

The 12 most recent posts are shown in full above; 20 posts from the same window are listed here in short form, each with its source post.

How to read this page

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest sourced?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a row to reach the sourced tier?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑