DiggTop ⌕

X local-model bench digest — Sep 26, 2026

As of Sep 26, 2026 — 12 posts collected from X. 32 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 26, 2026 07:05 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 15 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

M5 Ultra × Qwen3.8-Flash-Next

Sep 26, 2026 · by @Beamsters1

Context Quant Engine Decode Prefill
— — — — 3300.0 tok/s
— — — 240.0 tok/s —

M5 Ultra - Qwen3.8-Flash-Next, Prefill 3300 tok/s, Gen 240 tok/s Now we're talking right?

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Ryzen AI Max+ 395 with Radeon 8060S × Qwen 3.6 35B-A3B

Sep 26, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
— Q8 vulkan 53.6 tok/s —
— IQ1_M vulkan 20.5 tok/s —

i'm very impressed with my framework desktop, amd's strix halo with 128gb of unified memory. it's the workstation where i make my content and where my agents cut my videos. ryzen ai max+ 395 with a radeon 8060s, 128gb unified qwen 3.6 35b-a3b q8 at 53.6 tok/s on vulkan, no spec decode ran a 397b moe in 90gb (iq1_m) at 20.5 tok/s on vulkan 1024x1024 images in about 20 seconds in comfyui on native rocm 10 days of…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3060 × Qwen3.8-35B-A3B

Sep 26, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 Q4_K_M llama.cpp — 550.0 tok/s
262144 Q4_K_M llama.cpp 50.0 tok/s —

Qwen3.8-35B-A3B running at around 50 tok/s on an RTX 3060 12GB is already impressive. But the token speed isn’t what catches my attention. It’s the fact that we’re talking about a 35B model running on a GPU that most people would consider too small for the job. The setup is surprisingly modest: RTX 3060 12GB 16GB system RAM Q4_K_M 262K context ~550 tok/s prefill ~50 tok/s decode The model is Empero’s…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

AMD MI355X x512 × gpt-oss-120b

Sep 26, 2026 · by @mylifcc

每秒 5,750,000 个 Token! Crusoe 刚刚公布了最新的 MLPerf Inference v6.1 成绩单,用 512 张 AMD MI355X 组网,创下了算力集群的新纪录。 划重点: 🔹 gpt-oss-120b:575 万 tok/s(单卡 11.2k) 🔹 DeepSeek-R1:290 万 tok/s(单卡 5.7k) 几个非常亮眼的架构细节👇

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state).

RTX 3090 x2 × Qwen3.8 Flash

Sep 26, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— 3.5bpw exllamav2 105.0 tok/s —
— 3.5bpw exllamav2 92.0 tok/s —
— 3.5bpw exllamav2 105.0 tok/s —
— 3.5bpw exllamav2 120.0 tok/s —
— 3.5bpw exllamav2 — 1635.0 tok/s

Qwen3.8 Flash is starting to look very interesting on a pair of RTX 3090s. After testing different quantization levels, 3.5bpw looks like a pretty sweet spot between model quality and inference speed. The setup: • Qwen3.8 Flash • 2× RTX 3090 • 3.5bpw • ExLlamaV3 backend And the reported speeds are already pretty wild for a local setup. Around 105 tokens/s on average, with performance changing depending on what the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

M4 Max 128GB × Qwen3.8-Flash-Next

Sep 26, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— — mlx 63.0 tok/s —

Local run, no cloud. Qwen3.8-Flash-Next oQ4e 4-bit MLX (70GB resident) on an M4 Max 128GB via oMLX, MTP depth 6 63 tok/s on the speed suite. Thinking off vs on, same prompt: 5,120 tok / 103s / 1 console error 26,698 tok / 618s / 0 errors Thinking on also redesigned the whole composition. "Create a single self-contained HTML file using three.js from a CDN that renders an animated scene: a large rotating icosahedron…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, measurement method).

M4 Max × Qwen3.8-Flash-Next

Sep 26, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— — mlx 63.4 tok/s —
1491 — mlx 68.8 tok/s —
1499 — mlx 52.2 tok/s —
5938 — mlx 69.1 tok/s —
— NVFP4 vllm 48.5 tok/s —
1492 NVFP4 vllm 52.0 tok/s —
1270 NVFP4 vllm 44.3 tok/s —
667 NVFP4 vllm 49.2 tok/s —

Qwen3.8-Flash-Next Mac Studio vs single DGX Spark Mac Studio M4 Max (oQ4e 4-bit MLX, oMLX, native MTP depth 6) - mean 63.38 tok/s - sequence: 68.85 tok/s (1491 tok, TTFT 2.48s) - code: 52.20 tok/s (1499 tok, TTFT 0.78s) - json: 69.09 tok/s (5938 tok, TTFT 0.80s) DGX Spark ×1 (NVFP4, vLLM TP1, PLE offload) - mean 48.50 tok/s - sequence: 52.00 tok/s (1492 tok, TTFT 1.25s) - code: 44.28 tok/s (1270 tok, TTFT 1.02s) -…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 3060 × Bonsai 2 27B PTQ1_0

Sep 26, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
— — llama.cpp 50.0 tok/s —
131K Q4_K_M llama-server 41.1 tok/s —
131K Q4_K_M llama-server 65.8 tok/s —

from a 12gb gaming card to 2x dgx spark, here's the model i'd run at each tier, the speed i measured and the settings that make it work, bookmark this: 12gb: bonsai 2 27b ptq1_0 with the mtp head, 50 tok/s fresh on a 3060. kv cache at q4_0, one slot, n-max 1 on the head and reasoning effort medium, because the default xhigh can think through your whole token budget and hand you an empty answer. 24gb: qwen 3.8 27b…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

Unspecified platform × Qwen3.8-27B

Sep 25, 2026 · by @0xBakeer · third-party report

Context Quant Engine Decode Prefill
— NVFP4 — 50.0 tok/s —

Qwen3.8-27B in NVFP4 or Qwen3.8-flash-next in EXL3 at 2-bit. Both at 50 tok/s. Which one would you run?

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark x4 + RTX 5090 × MiMo-V2.6-Flash

Sep 25, 2026 · by @dangerm00se

Context Quant Engine Decode Prefill
— — deepseek rust engine 110.0 tok/s —
— — deepseek rust engine — 5105.0 tok/s

MiMo-V2.6-Flash, 4× DGX Spark + RTX 5090 hybrid: 110 tok/s decode (+54%) and 5,105 tok/s prefill (+72%) over vLLM on the Sparks alone. Opus 5.5 built it on @wrldsuksgo2mars's excellent DeepSeek Rust engine, with clever tricks from @Tech2Wild and @majewskizby. Links below 👇 https://t.co/DrOeKz23aD

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × GLM-5.3 Flash

Sep 25, 2026 · by @RunAnywhereAI

Context Quant Engine Decode Prefill
— — wally 363.0 tok/s —
— — flashx 143.0 tok/s —

We raced GLM-5.3 Flash on Wally against @Zai_org 's own paid fast tier, FlashX. Same prompt, live. Wally: 363 tok/s. FlashX: 143 tok/s. 2.5x faster, 3.5x cheaper. Why are you still using slow AI? Try Wally, $5 free. https://t.co/tku253unFd

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX PRO 6000 × MiMo-V2.6-Flash-RL

Sep 25, 2026 · by @ViC305

Context Quant Engine Decode Prefill
— 2.20 bpw SAGE-EXL3 exl3 184.1 tok/s —
— 2.20 bpw SAGE-EXL3 exl3 49.6 tok/s —
— 2.20 bpw SAGE-EXL3 exl3 — 2258.0 tok/s
— 2.20 bpw SAGE-EXL3 exl3 321.9 tok/s —

RTX PRO 6000 96GB + DGX Spark owners rejoice! 🔥 You can now run the highest quality Xiaomi’s MiMo-V2.6-Flash-RL locally in EXL3 on ONE DGX Spark or RTX 6000 that was done via my SAGE-EXL3 dynamic quantization process 🚀 309B total parameters. Only ~15B active per token. Xiaomi’s published agent results repeatedly place it in the same neighborhood as GPT-5.6 Sol and Claude Opus 5. And on ONE RTX PRO 6000: 184.1 tok/s…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑