DiggTop ⌕

X local-model bench digest — Sep 28, 2026

As of Sep 28, 2026 — 12 posts collected from X. 41 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 28, 2026 07:34 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 5 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x2 × Qwen3.8-Flash-Next-hibrid48

Sep 28, 2026 · by @vr8vr8

Context Quant Engine Decode Prefill
— NVFP4 vllm 0.30 99.0 tok/s —
— NVFP4 vllm 0.30 793.0 tok/s —
1K NVFP4 vllm 0.30 99.0 tok/s 2543.0 tok/s
128K NVFP4 vllm 0.30 89.8 tok/s 3233.0 tok/s
256K NVFP4 vllm 0.30 81.6 tok/s 3094.0 tok/s
— NVFP4 vllm 0.30 80.0 tok/s —

Qwen3.8-Flash-Next - 4.1v🥳 🚀 Peak c=1 @ 118 tok/s 📊 Average c=1 @ 80 tok/s 🏭 Peak c=64 @ 834 tok/s ⚙️ Prefill 3,233 tok/s 💻 TTFT 0.42s ⚡️ https://github.com/myllmbox/qwen38-flash-next-cluster-recipe https://t.co/JpR5pYH4Ee

Full methodology stated (engine, quantization, context, hardware state, method) — 3 added to the benchmark board (3 of 6 measured rows queued; the rest are published here only).

M5 Ultra 80 core GPU × Qwen 3.8 Flash Next

Sep 28, 2026 · by @iam_agg

Context Quant Engine Decode Prefill
32768 Q4 tensorfold 0.3.4.1 152.0 tok/s 2086.0 tok/s
32768 Q4 mlx 26.9.6 124.0 tok/s 3863.0 tok/s
32768 Q4 mlx 0.7.0rc1 105.0 tok/s 3694.0 tok/s

Rough M5 Ultra 80 core GPU 256GB benches: Qwen 3.8 Flash Next, Q4 variants, MTP, 1 stream, 32k context, averaged code/prose: TensorFold 0.3.4.1: 152 tok/s decode, 2086 tok/s prefill mlx-serve 26.9.6: 124 tok/s, 3863 tok/s prefill oMLX 0.7.0rc1: 105 tok/s, 3694 tok/s prefill https://t.co/ZhSKeUYxWa

Digest-only: incomplete methodology (hardware state).

RTX 3060 × Qwen3.8-35B-A3B

Sep 28, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 Q4_K_M llama.cpp — 550.0 tok/s
262144 Q4_K_M llama.cpp 50.0 tok/s —

Qwen3.8-35B-A3B running at around 50 tok/s on an RTX 3060 12GB is the part of local AI that I find genuinely interesting. Not because 50 tok/s is some ridiculous inference record. Because this is a 35B model running on a GPU that only has 12GB of VRAM. The setup is surprisingly ordinary: RTX 3060 12GB 16GB system RAM Q4_K_M 262K context ~550 tok/s prefill ~50 tok/s decode The model is Empero’s Qwen3.8-35B-A3B, a…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

M4 Max × Qwen3.8-27B

Sep 28, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — tensorfold 0.3.4 118.3 tok/s —
— — tensorfold 0.3.0 28.5 tok/s —
— — mlx MTP3 65.8 tok/s —
31K — tensorfold 0.3.4 69.0 tok/s —

TensorFold 0.3.4 just shipped with M1–M4 support, and the M4 Max numbers are kind of ridiculous. Earlier this morning, Qwen3.8-27B on the same Mac Studio M4 Max was doing 28.5 tok/s with TensorFold 0.3.0. Tonight, TensorFold 0.3.4 is hitting: 118.3 tok/s Same Qwen3.8-27B. Same M4 Max. Same three prompts. Thinking off. Median of five runs. That’s a 4.0x jump from the drafting path. The serial floor is around 29.3…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization).

RTX 5070 x1 × Qwen3.8-Flash-Next

Sep 28, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
128K Q2_0 strata 65.1 tok/s —
128K Q2_0 strata — 543.0 tok/s
128K IQ2_XS strata 52.0 tok/s —
128K IQ2_XS strata — 472.0 tok/s
128K IQ3_XXS strata 44.8 tok/s —
128K IQ3_XXS strata — 414.0 tok/s
— — llama.cpp 15.0 tok/s —

Qwen3.8-Flash-Next was running at around 15 tok/s on a 12GB RTX 5070. Apparently, that wasn’t good enough. So instead of swapping the GPU or giving up on the model, he built his own inference engine around it. That’s how Strata came about. And the reported jump is pretty ridiculous: llama.cpp: ~15 tok/s Strata: up to 65.1 tok/s That’s more than 4x the decode speed on the same machine. The hardware isn’t some massive…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

RTX 5060 Ti × Qwen3.8-27B

Sep 28, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262K — llama.cpp v1.1 67.0 tok/s —
39K — llama.cpp v1.1 39.5 tok/s —
119K — llama.cpp v1.1 21.0 tok/s —
262K — llama.cpp v1.1 53.0 tok/s —
262K — llama.cpp v1.1 42.0 tok/s —
262K — llama.cpp v1.1 71.0 tok/s —
262K — llama.cpp v1.1 66.0 tok/s —
262K — llama.cpp v1.1 127.0 tok/s —
119K — llama.cpp v1.1 — 1060.0 tok/s

Qwen3.8-27B is now doing something wild on a single RTX 5060 Ti 16GB. Bonsai 2, PrismML’s roughly 1.75-bit compression of Qwen3.8-27B, can run with: 262K context MTP head Vision All loaded simultaneously And the whole setup fits in about 15.2GB of the card’s 15.9GB usable VRAM. The reported speed is around 67 tok/s with the MTP head. For comparison: 67 tok/s with MTP 71 tok/s for code 53 tok/s without MTP 42 tok/s…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (measurement method).

DGX Spark × Qwen3.8-27B

Sep 28, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— NVFP4 custom inference engine 45.0 tok/s —

Qwen3.8-27B NVFP4 is now hitting 45 tok/s on a single DGX Spark, running on a custom inference engine. And the part I care about most isn’t just the speed. It’s that the output is lossless. Every token produced by the accelerated path matches what plain decoding would have produced. So this isn’t a case of making the model look faster by changing the output or using a different decoding path. The model is still…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark x3 × GLM-5.3 Flash

Sep 28, 2026 · by @jakeharrisdev

Context Quant Engine Decode Prefill
— — jspark3 v1.8 73.3 tok/s —
— — jspark3 v1.8 141.7 tok/s —
— — jspark3 v1.8 — 1262.6 tok/s

JSpark3 v1.8. GLM-5.3 Flash on three DGX Sparks, stock weights. 73.3 tok/s code on a single stream. 141.7 across 4 streams. 1,262.6 tok/s prefill. Everything's open: https://jakejh.com/jspark3/glm/?v=1.8 https://t.co/cdrrX73EvU

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).

Unspecified platform × GLM-5.3 Flash

Sep 27, 2026 · by @jakeharrisdev

Context Quant Engine Decode Prefill
— — jspark3 v1.8 73.3 tok/s —
— — jspark3 v1.8 49.1 tok/s —

GLM-5.3 Flash on JSpark3 v1.8 hits 73.3 tok/s on code and 49.1 tok/s on prose. Stock weights, pinned recipe, every number measured. Recipe and card on Hugging Face: https://huggingface.co/jakejharris/jspark3

Digest-only: hardware.name missing; incomplete methodology (context length, hardware state).

DGX Spark x3 × GLM-5.3 Flash

Sep 27, 2026 · by @jakeharrisdev

Context Quant Engine Decode Prefill
— — jspark3 v1.8 141.7 tok/s —
— — jspark3 v1.8 90.7 tok/s —

JSpark3 v1.8 is live. A big jump from v1.1, built with my agent fleet over the past few weeks. GLM-5.3 Flash. Up to 141.7 tok/s code and 90.7 tok/s prose decode at 4 streams on three DGX Sparks, stock weights. Pinned recipe, measured results. https://www.jakejh.com/jspark3/glm/?v=1.8

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state).

GB10 × GLM-5.3-Flash-EXL3-2.25bpw-sm121

Sep 27, 2026 · by @mr_r0b0t

Context Quant Engine Decode Prefill
8192 EXL3 2.25bpw + DFlash2 3.00bpw exllamav2 native 50.0 tok/s 22.6 tok/s

r0b0tlab/GLM-5.3-Flash-EXL3-2.25bpw-sm121 + 3.00bpw DFlash2 drafter + native ExLlamaV3 runtime! Tested on a single @NVIDIAAI GB10 ♥️ At 8K C1, code: 22.61→49.96 tok/s with K=5 Exact single-key retrieval at 259,993 prompt tokens Extended eval suite in progress 🤓 https://t.co/oYxwpC4I7S

Digest-only: incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark × Qwen3.8-27B

Sep 27, 2026 · by @0xBakeer

Context Quant Engine Decode Prefill
— NVFP4 my own engine 45.0 tok/s —

Qwen3.8-27B NVFP4 on one DGX Spark, running on my own engine. 45 tok/s on prose, and it's lossless: every token is exactly what plain decoding would write. All the other numbers are live in this 2 min demo. Still a WIP 🚧 I'll publish it this week. Really proud of this one https://t.co/OGRFYhoh35

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑