DiggTop ⌕

X local-model bench digest — Sep 27, 2026

As of Sep 27, 2026 — 12 posts collected from X. 32 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 27, 2026 09:30 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 14 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

GPU

Sep 27, 2026 · by @DailyDoseOfDS_

Context Quant Engine Decode Prefill
— — freetoken 39.3 tok/s —
— — freetoken 22.0 tok/s —
— — freetoken 14.9 tok/s —

UC Berkeley open-sourced FreeToken. (2–4x faster local LLM inference than Ollama) the results are wild: - Qwen3.6-35B on an 8GB GPU at 39.3 tokens/s - DeepSeek-V4-Flash 284B on a 32GB GPU at 22 tokens/s - GLM-5.2 753B on a 96GB GPU at 14.9 tokens/s a 35B model at 16-bit precision needs about 70GB just for its weights. even at 4 bits it is close to 18GB, and FreeToken serves it on an 8GB GPU. let us explain how: all…

Digest-only: model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash

Sep 27, 2026 · by @sfxnz

Context Quant Engine Decode Prefill
— EXL3 — 43.0 tok/s —
— EXL3 — 84.0 tok/s —
8K EXL3 — — 814.0 tok/s

Week long optimisations for Deepseek v4.1 flash EXL3 2X DGX sparks just merged⚡️ 30+ boots on my own sparks Speed, single stream, same hardware Long form prose: 32 → 43 tok/s (+31%) Highly predictable output (counting): 59 → 84 tok/s (+44%) Cold 8k token prefill: 184 →814 tok/s (4.4x) (fixed) https://t.co/KzBlJaWGf8

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

RTX 3090 × qwen

Sep 27, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
— — llama.cpp 41.0 tok/s —
— — llama.cpp 60.0 tok/s —
— — llama.cpp 191.9 tok/s —
— — llama.cpp 84.4 tok/s —

anyone who owns a rtx 3090 or any 24gb card and want qwen 3.8 27b dense running properly, or you just want to know what your gpu, or two of them, actually does with it, this is the repo i'd send you to, link below. it started as one llama.cpp flag. the mtp head, switching it on took a 3090 from about 41 tok/s to the mid 60 tok/s with nothing else changed. then strangers started running it on hardware i couldn't buy…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

EVO-X2 128GB × Qwen3.8-Flash-Next

Sep 27, 2026 · by @UtaAoya

Context Quant Engine Decode Prefill
256K — halogen 0.14.0 42.0 tok/s —
— — halogen 0.14.0 74.0 tok/s —
— — halogen 0.14.0 — 1230.0 tok/s

Local LLM - Qwen3.8-Flash-Next with EVO-X2 ”llama-split-bench方式ではない”ので、載せてもらえるかわかりませんが、PR出しました!☺️ 使用したのは”halogen”。Prifileの速さとDecodeが長コンテキストでも低下が小さいのが特徴です。これとV100 32GB x 3 環境を比較する予定 https://t.co/xhNLVmUFPs ◆チャッピー解説 EVO-X2 128GBで Qwen3.8-Flash-Next W4B を Halogen 0.14.0 で検証しました。 小型PC 1台でも、256K級の長Contextで実運用目安 約41〜43 tok/s を維持。Synthetic条件では 72〜76 tok/s、32Kでは MTPで約28.8%高速化 も確認できました。 さらに増分Prefillも 約1.16k〜1.30k tok/s…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (hardware state, measurement method).

Mac Studio M4 Max × Qwen3.8-27B

Sep 27, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— 4-bit tensorfold 0.3.4 76.7 tok/s —
— 4-bit tensorfold 0.3.4 57.4 tok/s —

Qwen-3.8-27b on Mac Studio m4 max, amazing results with thinking on, but off isn't bad eiter Same prompt. Same seed. Thinking on vs off Top: thinking off: 6,071 tokens, 82 s, 76.7 tok/s. Needed 2 repairs to run (typo'd CDN URL, 6 missing circle centers) Bottom: thinking on: 11,895 tokens, 25K chars of reasoning, 210 s, 57.4 tok/s. Zero repairs. Ran first try, correct 19‑circle Flower of Life, lattice. Qwen3.8‑27B…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).

M5 Ultra × GLM 5.3 Code

Sep 27, 2026 · by @JamesMcPherson

Context Quant Engine Decode Prefill
— — mlx 57.0 tok/s —
— — mlx 87.0 tok/s —
— — mlx 125.0 tok/s —

First "Big" finding on the M5 Ultra. TLDR: GLM 5.3 Code is running at 125 tok/s and prose at 87! Earlier in the day, I optimized oMLX for GLM5.3 and peaked at a very respectable 57 tok/s prose which is competitive with my 4x Spark cluster. After running Tensorflow and some tweaking for M5, (comments added to the PR for GLM5.3 so you can reproduce) I'm up to 87 tok/s prose and 125+ for code. Now, this is not really…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 3060 × MiMo-V2.6-Distill-Qwen-9B

Sep 27, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262K Q5_K_M llama.cpp 47.0 tok/s —
262K Q5_K_M llama.cpp — 1600.0 tok/s

Xiaomi’s MiMo-V2.6-Distill-Qwen-9B is running surprisingly comfortably on a single RTX 3060. And we’re talking about the 12GB version. The setup: RTX 3060 12GB Q5_K_M quantization 262K context ~47 tok/s decode ~1,600 tok/s prefill Getting a 9B model running locally isn’t unusual. Getting a full 262K context on a 12GB consumer GPU while still pushing around 47 tok/s is the part that caught my attention. The model…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Unspecified platform × llama.cpp

Sep 27, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
262K — ollama 50.0 tok/s —

if you're getting into local ai, the first mistake is taking the easy route. ollama and lm studio are good apps, but they hide exactly the parts you're supposed to learn, so you end up running models for months without knowing why they're fast, slow or out of memory. take the llama.cpp route instead. build it yourself, download one gguf, start llama-server and read the log. every line in it tells you something about…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

M4 Max × Qwen3.8-27B

Sep 27, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— — tensorfold 0.3.4 118.3 tok/s —
31K — tensorfold 0.3.4 — 227.9 tok/s
— — tensorfold 0.3.0 28.5 tok/s —

TensorFold 0.3.4 shipped tonight with M1–M4 support after our M4 Max result this morning. @ashxhart moved fast: 118tok/s Mac Studio M4 Max, Qwen3.8‑27B + DFlash2, same 3 prompts, thinking off, median of 5: 0.3.0 this morning: 28.5 tok/s 0.3.4 tonight: 118.3 tok/s serial floor 29.3 → 4.0× from drafting Byte‑identical to serial on all three. oMLX MTP3 on the same Mac: 65.8. The M4 Max now edges one DGX Spark on the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (hardware state).

Unspecified platform × Qwen3.8-Flash-Next

Sep 26, 2026 · by @seanhighness

Context Quant Engine Decode Prefill
190K — — 57.1 tok/s —
190K — — — 627.5 tok/s
— — — 55.0 tok/s —
— — — 41.0 tok/s —

PREFIX Caching was broken NOw fixed Qwen3.8-Flash-Next Working with vision Live: 192K context, MCS308 (308 CPU / 204 GPU experts), MTP3, vision and prefix reuse enabled. Determinism: cold and cached 256-token outputs matched at both 32K and 190K input lengths. 190K decode: 55.6–58.7 tok/s. 190K prefill: 161.1 seconds cold → 0.30 seconds cached. Quick https://t.co/49rEUmxXeQ: 55.0 tok/s code, 41.0 narrative. There's…

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX PRO 6000 Blackwell Max-Q x4 × DeepSeek-V4.1-Flash

Sep 26, 2026 · by @AiXsatoshi

Context Quant Engine Decode Prefill
— — vllm 187.2 tok/s —

510GBのDeepSeek-V4.1-Flash、 ローカルで動作確認しました Engine:vLLM GPU:RTX PRO 6000 Blackwell Max-Q ×4 構成:TP4、Engram RAM(DDR4)常駐 Single decode速度:約187.2 tok/s DSpark acceptance:92.2% ピークGPU電力:630 W(PL250W) 最大VRAM:92,278 MiB/GPU https://t.co/GV75JV9IYi

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, measurement method).

M3 Ultra × GLM-5.3-Flash

Sep 26, 2026 · by @volatilemarkts

Context Quant Engine Decode Prefill
— — mlx 28.0 tok/s —
— — tensorfold 60.0 tok/s —
— — tensorfold 79.0 tok/s —

I've been quietly worried for a year that my M3 Ultras would always be second-class citizens next to my DGX Sparks. Then @ashhart shipped TensorFold, and my M3 Ultra started acting like the M5 I have not bought yet. From 25 tok/s to 60tok/s... Thats not a typo. Same Mac. Same model. Same weights. Different engine: • GLM-5.3-Flash on the M3 Ultra: ~25–31 tok/s stock → 60 tok/s. First token lands before you finish…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑