DiggTop ⌕

X local-model bench digest — Oct 6, 2026

As of Oct 6, 2026 — 12 posts collected from X. 34 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 6, 2026 11:26 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 22 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x2 × GLM-5.3-Flash

Oct 6, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
131K NVFP4 — 25.0 tok/s —
262K — — 34.8 tok/s —
262K FP8 — 42.5 tok/s —

the results of 4 models compete the same floating tree build, i promised you this side by side when glm finished, here it is, ranked 1. glm 5.3 flash, the crown 👑 local on 2x dgx spark, 320b moe with 18b active, nvidia's nvfp4 20 to 30 tok/s with mtp k=3, 131k context, hermes agent 12/12 tests at 1h34m, 2,356 lines, visual gate at 32.6% canopy the fullest world from every angle, a dome of leaves and a goblin…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Oct 6, 2026 · by @jurlycat

Context Quant Engine Decode Prefill
100K NVFP4 dflash2 — 1500.0 tok/s
100K NVFP4 dflash2 40.0 tok/s —

Someone vibe-coded a GTA × Mirror’s Edge-style parkour game with a fully local model. GLM-5.3-Flash running on 2× DGX Spark, using NVFP4 quantization + DFlash2 speculative decoding. Claude Code as the coding harness. The builder reports: • ~1,500 tok/s prefill • ~40 tok/s decode at 100K context They debugged building collisions, then added flips, rolls and ledge grabs. This is what local AI progress looks like…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x1 × RED-SNOW-5.3-FLASH

Oct 6, 2026 · by @ViC305

Context Quant Engine Decode Prefill
— EXL3 2.49bpw exllamav2 38.7 tok/s —
— EXL3 2.49bpw tensorfold 41.0 tok/s —
— EXL3 2.49bpw rexl3 61.7 tok/s —

RED-SNOW-5.3-FLASH EXL3 is LIVE for both ONE and TWO DGX Sparks. 🥵❄️ ONE Spark: 2.49 bpw, 90.32% top-1, up to 61.7 tok/s TWO Sparks: 4.91 bpw, 96.85% top-1, KL 0.01174 New collab with @Blackfrost_AI. RED-SNOW is an abliteration of GLM-5.3-Flash with a cyber/security fine-tune, built for authorized red-team operations. I took the BF16 release and built two SAGE MixedK EXL3 targets depending on how much hardware you…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark × GLM-5.3-Flash

Oct 6, 2026 · by @NFT_Chen

Context Quant Engine Decode Prefill
262K EXL3 2.05bpw tensorfold v0.6.5 32.0 tok/s —
80K EXL3 2.05bpw tensorfold v0.6.5 28.5 tok/s —
250K EXL3 2.05bpw tensorfold v0.6.5 26.0 tok/s —

💥什么魔法?!320B 的 GLM-5.3-Flash 在一台 Spark 跑到 32 tok/s ! TensorFold 把同一套 EXL3 2.05bpw 从 11.63 拉到 32.32 tok/s(2.78x),单机吃下 262k 上下文,11k 对话的下一轮,从每次重灌 47.6 秒,变成复用 prompt state 的 2.1 秒,Agent 不用每步重读整段对话了。 主要亮点 🔹单卡 GB10,不是双机 CX7 🔹长 prompt 仍稳:8 万 token 约 28.5 tok/s,25 万仍有 26 tok/s 🔹下一轮首 token 2.1 秒,回复与重灌路径一致 🔹OpenAI 兼容接口,thinking 有上限,防空转 关键步骤 1️⃣clone TensorFold,切到 v0.6.5 的钉死 commit 2️⃣按序打两张补丁:单卡 EXL3 加载 + serving / prompt-state…

Digest-only: incomplete methodology (hardware state, measurement method).

Unspecified platform × TP5 GLM 5.3

Oct 6, 2026 · by @BrandonMusicKy

Context Quant Engine Decode Prefill
— — — — 2000.0 tok/s
— — — 76.0 tok/s —
— — — 272.0 tok/s —

tp5 glm 5.3 (full version) is working. 1.5 Million KV cache. 10/10 on estonia, 9/10 on lavd (high reasoning) (5 exact 4 near, 1 miss). prefill over 2,000 t/s. 76 tok/s single-stream (98–113 t/s on long reasoning) and 272 tok/s at 8 concurrent requests. not optimized but runs well.

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform

Oct 6, 2026 · by @cholf5 · third-party report

Context Quant Engine Decode Prefill
— — — 30.0 tok/s —

AI 用 React 做游戏真是神一样的存在,像这种界面,给个原版截图,40 分钟就搞成这样了,UI 没用到暴雪的资源,全是它用 css 自己做的。这还是 GLM Flash 这种 30 tok/s 的模型,要是 DS 估计几分钟就做好了。所以说选择大于努力,选 Unity 或 Godot,你就跟 Prefab 和结点树斗去吧。 https://t.co/HgePvkqmEo

Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × Qwen3.8-27B

Oct 6, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — custom llm engine 273.0 tok/s —

Qwen3.8-27B is hitting 273 tok/s decode on a custom LLM engine. And that’s probably more interesting than the number itself. The model is running with the TQ3_4S-v2 quantization, but the real work is happening underneath it. This is a custom inference engine being built specifically to squeeze more performance out of local models. For context, the same Qwen3.8-27B TQ3_4S-v2 ecosystem has already shown how much the…

Digest-only: hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3090 x1 × Qwen3.8-27B

Oct 6, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
25K — vllm 381.0 tok/s —
— — vllm 133.0 tok/s —
25K — vllm 260.0 tok/s —

Qwen3.8-27B just hit around 381 tok/s on a single RTX 3090. 24GB of VRAM. But the interesting part isn’t the 381 tok/s number. It’s how the inference engine got there. This isn’t a normal-chat benchmark where the model is freely generating 381 new tokens every second. The hardware stayed the same throughout: 1× RTX 3090 24GB VRAM ~250W The performance progression came almost entirely from improving the inference…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization).

Unspecified platform

Oct 6, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
656K — llama.cpp stock 42.1 tok/s —
— — custom kernel + speed head 50.1 tok/s —
96K — custom kernel + speed head 42.8 tok/s —
— — custom kernel + speed head 41.3 tok/s —

dig out the gaming gpu you retired, a 2016 gtx 1080 runs gemma 4 e4b at 42 tok/s and one of you in my replies is pushing a 27b at 50 tok/s on a 2020 rtx 3080 gtx 1080 8gb, from 2016: gemma 4 e4b at 42.13 tok/s with a 656k window on stock llama.cpp rtx 3060 12gb, from 2021: a 27b at 50.1 tok/s with my kernel and the speed head rtx 3060 ti 8gb: a 27b at 42.8 tok/s held all the way to a 96k window rtx 3090 24gb, from…

Digest-only: hardware.name missing; model.name missing; incomplete methodology (engine+version, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Oct 6, 2026 · by @yume_arasaki

Context Quant Engine Decode Prefill
192K Q3 strata 30.0 tok/s —
32K Q2_0 strata 94.0 tok/s 2650.0 tok/s
65K — llama.cpp 28.0 tok/s —
262K EXL3 2.5bpw llama.cpp 38.6 tok/s —
— IQ3_S strata 60.0 tok/s 2500.0 tok/s
— IQ3_S strata 160.0 tok/s —
262K — pentacoxian-dev 60.0 tok/s —

Insane developments have been happening on Qwen 3.8 Flash. In 2 weeks, suddenly everyones computer can run Qwen 3.8 Flash. I wanted to find out what happened and why everyone in local AI is using this model now. Every number below is a receipt I verified, my own hardware or community, and the ones from my fleet live in my bench repo. Here is what I found. WHY THIS MODEL Qwen3.8-Flash-Next is a 180B parameter MoE…

Digest-only: hardware.name missing; incomplete methodology (engine+version, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Oct 6, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — strata stack 62.0 tok/s —
— — strata stack — 2727.0 tok/s

Qwen3.8-Flash-Next just got another serious inference boost. In only four days, the Strata stack has pushed performance up by roughly 30% and finally crossed the 100 tok/s mark with TP=1. The latest numbers: 111 tok/s peak at concurrency 1 62 tok/s average at concurrency 1 414 tok/s peak at concurrency 16 2,727 tok/s prefill 0.57s TTFT That first number is the one I keep coming back to. 111 tok/s on a single…

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark × GLM-5.3 Flash

Oct 6, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— — — 26.0 tok/s —
— — — 32.0 tok/s —

Same koi pond prompt, two GLMs on DGX Sparks: -GLM-5.3 Flash (2 Sparks) 90,852 tokens · 59 min · 26 tok/s Worked first try, no repairs -Full GLM-5.3 753B (4 Sparks) 136,640 tokens · 70 min · 32 tok/s Bigger world, but needed a 7-line fix to show its water Bigger model, more thinking, not automatically a better first draft.

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑