DiggTop ⌕

X local-model bench digest — Sep 29, 2026

As of Sep 29, 2026 — 12 posts collected from X. 42 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 29, 2026 13:03 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 25 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark × Qwen3.8-27B

Sep 29, 2026 · by @ViC305

Context Quant Engine Decode Prefill
— EXL3 3.00 bpw tensorfold 0.3.6 83.4 tok/s —
— EXL3 3.00 bpw mlx 57.5 tok/s —
— EXL3 3.00 bpw vllm 23.4 tok/s —
— EXL3 3.05 bpw tensorfold 0.3.6 80.8 tok/s —
— EXL3 3.05 bpw vllm 42.4 tok/s —
— EXL3 3.05 bpw tensorfold 0.3.6 77.7 tok/s —
— EXL3 3.05 bpw vllm 40.9 tok/s —
— EXL3 3.05 bpw tensorfold 0.3.6 69.2 tok/s —
— EXL3 3.05 bpw vllm 37.6 tok/s —

TensorFold 0.3.6 just took CUDA EXL3 from a GLM-only experiment to a reusable mixed-bit backend. 🤯 I wrote the shared EXL3 module that made that possible.🚀 49 files. +5,673 / -37. 3 codebooks. 1–8 bits. Mixed precision per tensor. ✍️ And it SHIPPED. ✅ Before this, TensorFold’s CUDA EXL3 path was built specifically around GLM-5.3-Flash’s 4-bit MCG experts. Now the same infrastructure serves: → Qwen3.8-27B EXL3 +…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Sep 29, 2026 · by @jayleaton

Context Quant Engine Decode Prefill
— — vllm 42.6–44.0 68.6 tok/s —
— — vllm 18.9–22.2 43.2 tok/s —
— — — 88.2 tok/s —
— — — 82.2 tok/s —

Ran @alexellisuk's RigMark on my TensorFold build of GLM-5.3-Flash, 2x DGX Spark. Standard suite, settings untouched. code: 68.6 tok/s (vLLM 42.6–44.0) prose: 43.2 (vLLM 18.9–22.2) structured: 88.2 4 concurrent: 82.2 Lower than my own bench this morning, which I expected, but actually running that test pointed out some improvements I missed on my first pass. Bit behind on cold prefill and replay. This is the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state).

DGX Spark x4 × GLM-5.3-Flash

Sep 29, 2026 · by @majewskizby

Context Quant Engine Decode Prefill
— NVFP4 sparkdash 90.3 tok/s —
— NVFP4 sparkdash — 3500.0 tok/s
— NVFP4 sparkdash 113.8 tok/s —
— NVFP4 sparkdash 62.1 tok/s —
— NVFP4 sparkdash 153.2 tok/s —

Update: GLM-5.3-Flash TP4 on 4× DGX Spark 🚀 • Sparkdash prose decode: 90.3 tok/s at c1 (was 70, and ~37 on Sep 18) • Prefill: ~3.5k tok/s at 16k–128k (was ~2.2k) • RigMark: code 113.8 · prose 62.1 · structured 153.2 • Lossless: NVFP4 experts + 8-bit dense, draft verified token by token • GSM8K 98.4% · HumanEval 157/164 Recipe, measurements and credits: https://t.co/MbYYhBjlz8

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3090 × Qwen Flash Next

Sep 29, 2026 · by @needmorevram

Context Quant Engine Decode Prefill
128K — — 75.4 tok/s —

By removing vision, you can push the RTX 3090 even further. /I’m now seeing 75.4 tok/s decode on Qwen Flash Next, with 128K context, 23.8 / 24GB VRAM used, 8,182 experts cached, and an 82.6% hit rate. That’s on a single 3090 running at PCIe 3.0 x8. https://t.co/nYkYo3itHJ

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).

RTX 3090 × Qwen Flash Next

Sep 29, 2026 · by @needmorevram

Context Quant Engine Decode Prefill
128K IQ3_S GSQ RCO — 72.0 tok/s —

Back with another test on the new Strata release. Running IQ3_S GSQ RCO Qwen Flash Next at 128K context on a single RTX 3090, and I’m now seeing 72 tok/s decode. That is honestly mind blowing 🤯 https://t.co/vhLiZbiY6U

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

RTX 3090 × Qwen3.8-Flash-Next

Sep 29, 2026 · by @0xSero

Context Quant Engine Decode Prefill
3 EXL3 3.05bpw_h5_ng5 sglang v0.7.0-ampere 85.5 tok/s —

Qwen3.8-Next-Flash Unfortunately you'll need a total of 96gb system ram for it to work reliably, ends up at 85.5 tok/s at c3 1. https://github.com/0xSero/local-ai-images/pkgs/container/sglang-exl3-flashnext 2. https://github.com/0xSero/local-ai-images/pkgs/container/sglang-exl3-xpu-flashnext Enjoy! https://t.co/3WN0JJF4V1

Digest-only: incomplete methodology (measurement method).

DGX Spark x4 + RTX 5090 × GLM-5.3-Flash

Sep 29, 2026 · by @dangerm00se

Context Quant Engine Decode Prefill
— EXL3 deepseek rust engine 124.0 tok/s —
— EXL3 deepseek rust engine — 5300.0 tok/s

GLM-5.3-Flash on 4 DGX Sparks + a 5090, with @BrandonMusicKy's EXL3 weights: 124 tok/s decode, 5.3K tok/s prefill, well over our favourite 4-Spark recipes by @Tech2Wild and @MiaAI_lab. Continuing what tj (@wrldsuksgo2mars) started with his DeepSeek Rust engine. https://t.co/fGpB7f0FfR

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD

Sep 29, 2026 · by @0xSero

Context Quant Engine Decode Prefill
200K — — 65.0 tok/s —
200K — — — 2000.0 tok/s

You can now run an LLM which scores higher than Sonnet-5 medium, and GPT-6-Sol medium for 3000$ in compute - 65 decode tok/s - 2000 prefill tok/s - 200k kv cache 1. RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD 2. 64 GB of RAM 3. 100 GB of NVMe Recipe today https://t.co/DH0PZm7ROp

Digest-only: model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 3090 x4 × Qwen3.8 Flash Next

Sep 29, 2026 · by @superalesha

Context Quant Engine Decode Prefill
256K EXL3 4.05 exllamav2 90.0 tok/s —
256K EXL3 4.05 exllamav2 35.0 tok/s —

I tested 4 quant variants of Qwen3.8 Flash Next to find what works best on my 4x RTX 3090 rig. EXL3 4.05, 5.05 and 6.05 bpw from @turboderp_ , plus GSQ-RCO IQ3_S GGUF. 85 speed measurements across context lengths, KV settings, MTP and concurrent requests. All at 300W per GPU, with vision enabled. I wanted real 256K context, ideally for 2 sessions at once. My pick from these runs: EXL3 4.05 + FP16 GPU KV + MTP2. At…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 3090 × Qwen3.8-Next-Flash

Sep 29, 2026 · by @0xSero

Context Quant Engine Decode Prefill
— — — — 2000.0 tok/s
— — — — 2800.0 tok/s
— — — 64.0 tok/s —
— — — — 800.0 tok/s
— — — — 1200.0 tok/s
— — — 35.0 tok/s —

GPU poor rejoice! You will be able to run Qwen3.8-Next-Flash on a single 24-32 GB GPU + 64GB ddr4/ddr5 + 100GB NVMe/SSD RTX 3090: - 2000 / 2800 tok/s prefill - 64 tok/s decode c1 - 210k fp8 cache Intel Arc B70: - 800 / 1200 prefill tok/s - 35 decode tok/s - 270k fp8 cache https://t.co/Bh72r5qSaa

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x1 × Qwen3.8-Flash-Next

Sep 29, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
256K — tensorfold 62.0 tok/s —
256K — tensorfold 119.0 tok/s —
256K — tensorfold — 2500.0 tok/s

Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥 This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements! - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared to vLLM! Performance: Decode prose 62+ tok/s single stream Decode prose 119+ tok/s on 5 streams Prefill is mostly 2500 tok/s across the board!…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Sep 29, 2026 · by @jayleaton

Context Quant Engine Decode Prefill
— — tensorfold 44.6 tok/s —
— — vllm 22.8 tok/s —
— — tensorfold 77.6 tok/s —
— — vllm 41.9 tok/s —
— — tensorfold — 1607.0 tok/s
— — vllm — 1448.0 tok/s

GLM-5.3-Flash on 2x DGX Spark, same weights. TensorFold vs vLLM: chat 44.6 vs 22.8 tok/s code 77.6 vs 41.9 tok/s prompt processing 1,607 vs 1,448 tok/s Prompt processing was losing to vLLM yesterday (0.8x). Fixed it overnight. Found a few more spots with speed left on the table, so reworking that part now. Anyone else running GLM Flash on Sparks? Have you tried my repo yet?

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑