DiggTop

X local-model bench digest — Sep 22, 2026

As of Sep 22, 2026 — 12 posts collected from X. 41 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 22, 2026 11:05 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 17 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x2 × MiMo-V2.6-Flash

Sep 22, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
1M eagle mtp 35.0 tok/s
1M eagle mtp 46.0 tok/s
1M eagle mtp 120.0 tok/s
1M eagle mtp 138.0 tok/s
1M dflash 26.0 tok/s
1M dflash 70.0 tok/s
1M dflash 103.5 tok/s
1M dflash 190.0 tok/s

Run MiMo-V2.6-Flash on 2x DGX Sparks ⚡️ - Native 1M context, 2.9M KV cache pool - Full OMNI: text + image + video + audio Two different runtimes, different numbers. EAGLE MTP (default): - decode on prose single stream: 35 tok/s - decode on code single stream: 46 tok/s - decode on prose, 8 streams: 120 tok/s - decode on code, 8 streams: 138 tok/s DFlash draft model - 1.6M KV cache pool: - decode on prose: 26 tok/s -…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

M5 Max 64GB × Qwen3.8-Flash-Next

Sep 22, 2026 · by @Oluwaphilemon1 · third-party report

Context Quant Engine Decode Prefill
EXL3 hybrid mlx 70.0 tok/s

Qwen3.8-Flash-Next is getting interesting on Apple Silicon. A new EXL3 hybrid setup reportedly brings the model down to around 49GB of memory residency, which means an M5 Max Mac with 64GB unified memory can actually run it with room to spare. And the reported speed is 70+ tok/s. That’s a pretty different local AI experience for a model in this class. The interesting part isn’t simply that the model fits. It’s the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform

Sep 22, 2026 · by @Beamsters1

Context Quant Engine Decode Prefill
1M 3500.0 tok/s
1M 70.0 tok/s

I spent like 2-3 weeks working on this kernel, it achieved 3500 tok/s prefill and 70 tok/s decode at 1m context, but I think there's massive room to improve - only when I get my hand on the M5 Ultra unit. Not soon unfortunately.

Digest-only: hardware.name missing; model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Sep 22, 2026 · by @0xBakeer

Context Quant Engine Decode Prefill
EXL3 mtp 64.6 tok/s
EXL3 mtp 7.2 tok/s
EXL3 mtp 31.1 tok/s

I benchmarked the recipe of @ViC305 and I’m impressed. Qwen3.8-Flash-Next on EXL3 decodes at 64.6 tok/s for one user. The speed comes from MTP: the model drafts 5 tokens ahead and keeps the ones it would have written anyway. That trick does not survive a crowd. With 8 users at once, each one gets 7.2 tok/s. The whole server makes 31.1 tok/s, while vLLM on the same model makes 78.6. Quality holds. It scored 0.993 on…

Digest-only: hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3060 × MiMo-V2.6-Distill-Qwen-9B

Sep 22, 2026 · by @DogukanUrker

Context Quant Engine Decode Prefill
262144 Q5_K_M llama.cpp 47.0 tok/s
262144 Q5_K_M llama.cpp 1600.0 tok/s

running Xiaomi's MiMo-V2.6-Distill-Qwen-9B at 5-bit (Q5_K_M) on a single RTX 3060. 12GB VRAM. (config below) full 262K context. ~47 tok/s decode, ~1600 tok/s prefill. it's Qwen3.5-9B fine-tuned on MiMo-V2.6 outputs. Xiaomi reports SWE Pro going 32.0 - 44.6 over the base model. running my own agentic coding benchmark on it next! config: llama-server -m MiMo-V2.6-Distill-Qwen-9B-Q5_K_M.gguf -ngl 99 -c 262144 -fa on…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

MI300X × MiMo-V2.6-Pro

Sep 22, 2026 · by @tinygrad

Context Quant Engine Decode Prefill
tinygrad 200.0 tok/s

~200 tok/s MiMo-V2.6-Pro on MI300X brought up by GLM-5.3 https://t.co/VxXAZET5As

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

M5 Ultra × DeepSeek V4 Flash

Sep 22, 2026 · by @wquguru · third-party report

Context Quant Engine Decode Prefill
340.0 tok/s

同样一万刀跑本地 DeepSeek V4 Flash: 生成速度几乎打平 预填充 Spark 比 M5 Ultra 快 3–5 倍 并发 Spark 能堆到 340 tok/s,而 Ultra 顶多 2–3 并发 如果让我选的话,纯跑本地大模型,大概率还是不会选 M5 Ultra,论生态支持、可扩展性、综合性价比,Spark 还是遥遥领先。你们会买哪台?

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × Qwen3.8

Sep 22, 2026 · by @udonuopn_pc · third-party report

Context Quant Engine Decode Prefill
Q4_K_M 75.0 tok/s

5080追加するのとRAM増やすのと、どっちが速くなりそう? CPU:Ryzen 9950X RAM:32GB GPU:RTX5080 + RTX4080Super VRAM:16GB + 16GB = 32 GB Qwen3.8 27B Q4_K_M:70〜75 tok/s ?

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

1 gpu × Mimo 2.6 Flash

Sep 22, 2026 · by @SSHCodes

Context Quant Engine Decode Prefill
30.0 tok/s
  • get 1 gpu - try to run mimo 2.6 flash - get 30 tok/s decode speed, be excited - 30 tok/s PREFILL SPEED - realize you cannot run the goated model and that you need another gpu 😢

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × MiMo-V2.6-Flash

Sep 22, 2026 · by @Tech2Wild

Context Quant Engine Decode Prefill
300K vllm 88.0 tok/s
300K vllm 85.0 tok/s
300K vllm 70.0 tok/s
300K vllm 69.0 tok/s
300K vllm 57.0 tok/s
300K vllm 26.0 tok/s
300K vllm 159.0 tok/s
300K vllm 206.0 tok/s
300K vllm 249.0 tok/s
300K vllm 296.0 tok/s

REPO is UPDATED ! 🔥 MiMo-V2.6-Flash · 2x DGX Spark TP2 · vLLM + DFlash · image + video + audio Benched warmed + streaming, fixed 9-category prompt set, C1 to C6: ⚡️ 88 tok/s single-stream on counting · 85 structured tables · 70 code · 69 math · 57 JSON · 26 prose 🚀 159 tok/s aggregate at 6 streams · 206 on code · 249 on tables · 296 on counting 🧠 1.87M-token fp8 KV · 300K ctx serving · 6 full 300K requests at once ·…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state).

DGX Spark x2 × MiMo-V2.6-Flash

Sep 22, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
vllm 44.0 tok/s

MiMo-V2.6-Flash on 2× DGX Spark, day-zero, @Tech2Wild vLLM recipe. 44 tok/s Rocket-launch brief, thinking on, no cap: 80.3K tokens in 30 min: 0.5 s TTFT, DFlash accepting 4–5 tokens/step in the code phase. Complete file, first try, zero errors. Same prompt on GLM-5.3-Flash, same pair: 79.6K tokens, 57 min, at about 23 tok/s. Same thinking budget, half the wait. MiMo is way faster but GLM's video is a bit better…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark GB10 × Qwen3.8-Flash-Next

Sep 22, 2026 · by @ViC305

Context Quant Engine Decode Prefill
EXL3 3.05bpw exllamav2 78.3 tok/s
NVFP4 vllm 31.0 tok/s
EXL3 3.05bpw exllamav2 600.0 tok/s
NVFP4 vllm 150.0 tok/s
EXL3 3.05bpw exllamav2 569.0 tok/s
NVFP4 vllm 176.0 tok/s
EXL3 3.05bpw exllamav2 52.0 tok/s
NVFP4 vllm 64.8 tok/s
EXL3 3.05bpw exllamav2 66.4 tok/s
NVFP4 vllm 82.1 tok/s

Qwen3.8-Flash-Next: EXL3 ~600 vs NVFP4 ~150 tok/s reported concurrent decode. 🤯 Both on DGX Spark GB10. The interesting part is what the engines actually do with those compressed weights. This will be my last informational post on Qwen3.8-Flash-Next as im cooking up something new for you all! 👨🏻‍🍳 3.05 bpw EXL3 through my ExLlamaV3 fork. NVFP4 through vLLM. 𝗧𝗛𝗘 𝗗𝗘𝗖𝗢𝗗𝗘 𝗗𝗜𝗙𝗙𝗘𝗥𝗘𝗡𝗖𝗘 Balanced workload, EXL3 / NVFP4…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.