X local-model bench digest — Sep 29, 2026
As of Sep 29, 2026 — 12 posts collected from X. 42 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 29, 2026 13:03 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 25 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
DGX Spark × Qwen3.8-27B
Sep 29, 2026 · by @ViC305
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 3.00 bpw | tensorfold 0.3.6 | 83.4 tok/s | — |
| — | EXL3 3.00 bpw | mlx | 57.5 tok/s | — |
| — | EXL3 3.00 bpw | vllm | 23.4 tok/s | — |
| — | EXL3 3.05 bpw | tensorfold 0.3.6 | 80.8 tok/s | — |
| — | EXL3 3.05 bpw | vllm | 42.4 tok/s | — |
| — | EXL3 3.05 bpw | tensorfold 0.3.6 | 77.7 tok/s | — |
| — | EXL3 3.05 bpw | vllm | 40.9 tok/s | — |
| — | EXL3 3.05 bpw | tensorfold 0.3.6 | 69.2 tok/s | — |
| — | EXL3 3.05 bpw | vllm | 37.6 tok/s | — |
TensorFold 0.3.6 just took CUDA EXL3 from a GLM-only experiment to a reusable mixed-bit backend. 🤯 I wrote the shared EXL3 module that made that possible.🚀 49 files. +5,673 / -37. 3 codebooks. 1–8 bits. Mixed precision per tensor. ✍️ And it SHIPPED. ✅ Before this, TensorFold’s CUDA EXL3 path was built specifically around GLM-5.3-Flash’s 4-bit MCG experts. Now the same infrastructure serves: → Qwen3.8-27B EXL3 +…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).
DGX Spark x2 × GLM-5.3-Flash
Sep 29, 2026 · by @jayleaton
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | vllm 42.6–44.0 | 68.6 tok/s | — |
| — | — | vllm 18.9–22.2 | 43.2 tok/s | — |
| — | — | — | 88.2 tok/s | — |
| — | — | — | 82.2 tok/s | — |
Ran @alexellisuk's RigMark on my TensorFold build of GLM-5.3-Flash, 2x DGX Spark. Standard suite, settings untouched. code: 68.6 tok/s (vLLM 42.6–44.0) prose: 43.2 (vLLM 18.9–22.2) structured: 88.2 4 concurrent: 82.2 Lower than my own bench this morning, which I expected, but actually running that test pointed out some improvements I missed on my first pass. Bit behind on cold prefill and replay. This is the…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state).
DGX Spark x4 × GLM-5.3-Flash
Sep 29, 2026 · by @majewskizby
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | sparkdash | 90.3 tok/s | — |
| — | NVFP4 | sparkdash | — | 3500.0 tok/s |
| — | NVFP4 | sparkdash | 113.8 tok/s | — |
| — | NVFP4 | sparkdash | 62.1 tok/s | — |
| — | NVFP4 | sparkdash | 153.2 tok/s | — |
Update: GLM-5.3-Flash TP4 on 4× DGX Spark 🚀 • Sparkdash prose decode: 90.3 tok/s at c1 (was 70, and ~37 on Sep 18) • Prefill: ~3.5k tok/s at 16k–128k (was ~2.2k) • RigMark: code 113.8 · prose 62.1 · structured 153.2 • Lossless: NVFP4 experts + 8-bit dense, draft verified token by token • GSM8K 98.4% · HumanEval 157/164 Recipe, measurements and credits: https://t.co/MbYYhBjlz8
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 3090 × Qwen Flash Next
Sep 29, 2026 · by @needmorevram
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 128K | — | — | 75.4 tok/s | — |
By removing vision, you can push the RTX 3090 even further. /I’m now seeing 75.4 tok/s decode on Qwen Flash Next, with 128K context, 23.8 / 24GB VRAM used, 8,182 experts cached, and an 82.6% hit rate. That’s on a single 3090 running at PCIe 3.0 x8. https://t.co/nYkYo3itHJ
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
RTX 3090 × Qwen Flash Next
Sep 29, 2026 · by @needmorevram
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 128K | IQ3_S GSQ RCO | — | 72.0 tok/s | — |
Back with another test on the new Strata release. Running IQ3_S GSQ RCO Qwen Flash Next at 128K context on a single RTX 3090, and I’m now seeing 72 tok/s decode. That is honestly mind blowing 🤯 https://t.co/vhLiZbiY6U
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 3090 × Qwen3.8-Flash-Next
Sep 29, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 3 | EXL3 3.05bpw_h5_ng5 | sglang v0.7.0-ampere | 85.5 tok/s | — |
Qwen3.8-Next-Flash Unfortunately you'll need a total of 96gb system ram for it to work reliably, ends up at 85.5 tok/s at c3 1. https://github.com/0xSero/local-ai-images/pkgs/container/sglang-exl3-flashnext 2. https://github.com/0xSero/local-ai-images/pkgs/container/sglang-exl3-xpu-flashnext Enjoy! https://t.co/3WN0JJF4V1
Digest-only: incomplete methodology (measurement method).
DGX Spark x4 + RTX 5090 × GLM-5.3-Flash
Sep 29, 2026 · by @dangerm00se
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | deepseek rust engine | 124.0 tok/s | — |
| — | EXL3 | deepseek rust engine | — | 5300.0 tok/s |
GLM-5.3-Flash on 4 DGX Sparks + a 5090, with @BrandonMusicKy's EXL3 weights: 124 tok/s decode, 5.3K tok/s prefill, well over our favourite 4-Spark recipes by @Tech2Wild and @MiaAI_lab. Continuing what tj (@wrldsuksgo2mars) started with his DeepSeek Rust engine. https://t.co/fGpB7f0FfR
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD
Sep 29, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 200K | — | — | 65.0 tok/s | — |
| 200K | — | — | — | 2000.0 tok/s |
You can now run an LLM which scores higher than Sonnet-5 medium, and GPT-6-Sol medium for 3000$ in compute - 65 decode tok/s - 2000 prefill tok/s - 200k kv cache 1. RTX 3090 / 4090 / 5090 / Intel Arc B70 / AMD 2. 64 GB of RAM 3. 100 GB of NVMe Recipe today https://t.co/DH0PZm7ROp
Digest-only: model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
RTX 3090 x4 × Qwen3.8 Flash Next
Sep 29, 2026 · by @superalesha
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256K | EXL3 4.05 | exllamav2 | 90.0 tok/s | — |
| 256K | EXL3 4.05 | exllamav2 | 35.0 tok/s | — |
I tested 4 quant variants of Qwen3.8 Flash Next to find what works best on my 4x RTX 3090 rig. EXL3 4.05, 5.05 and 6.05 bpw from @turboderp_ , plus GSQ-RCO IQ3_S GGUF. 85 speed measurements across context lengths, KV settings, MTP and concurrent requests. All at 300W per GPU, with vision enabled. I wanted real 256K context, ideally for 2 sessions at once. My pick from these runs: EXL3 4.05 + FP16 GPU KV + MTP2. At…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
RTX 3090 × Qwen3.8-Next-Flash
Sep 29, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | — | 2000.0 tok/s |
| — | — | — | — | 2800.0 tok/s |
| — | — | — | 64.0 tok/s | — |
| — | — | — | — | 800.0 tok/s |
| — | — | — | — | 1200.0 tok/s |
| — | — | — | 35.0 tok/s | — |
GPU poor rejoice! You will be able to run Qwen3.8-Next-Flash on a single 24-32 GB GPU + 64GB ddr4/ddr5 + 100GB NVMe/SSD RTX 3090: - 2000 / 2800 tok/s prefill - 64 tok/s decode c1 - 210k fp8 cache Intel Arc B70: - 800 / 1200 prefill tok/s - 35 decode tok/s - 270k fp8 cache https://t.co/Bh72r5qSaa
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x1 × Qwen3.8-Flash-Next
Sep 29, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256K | — | tensorfold | 62.0 tok/s | — |
| 256K | — | tensorfold | 119.0 tok/s | — |
| 256K | — | tensorfold | — | 2500.0 tok/s |
Qwen3.8-Flash-Next for a single DGX Spark got a serious upgrade with TensorFold🔥 This is a completely new recipe, optimized and tuned for TesnorFold! Expect further improvements! - KV cache pool is ~1.3M - Default context 256k, with 5 concurrent. - Faster everything compared to vLLM! Performance: Decode prose 62+ tok/s single stream Decode prose 119+ tok/s on 5 streams Prefill is mostly 2500 tok/s across the board!…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).
DGX Spark x2 × GLM-5.3-Flash
Sep 29, 2026 · by @jayleaton
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | tensorfold | 44.6 tok/s | — |
| — | — | vllm | 22.8 tok/s | — |
| — | — | tensorfold | 77.6 tok/s | — |
| — | — | vllm | 41.9 tok/s | — |
| — | — | tensorfold | — | 1607.0 tok/s |
| — | — | vllm | — | 1448.0 tok/s |
GLM-5.3-Flash on 2x DGX Spark, same weights. TensorFold vs vLLM: chat 44.6 vs 22.8 tok/s code 77.6 vs 41.9 tok/s prompt processing 1,607 vs 1,448 tok/s Prompt processing was losing to vLLM yesterday (0.8x). Fixed it overnight. Found a few more spots with speed left on the table, so reworking that part now. Anyone else running GLM Flash on Sparks? Have you tried my repo yet?
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @ViC305 — Sep 29, 2026
- @jayleaton — Sep 29, 2026
- @majewskizby — Sep 29, 2026
- @needmorevram — Sep 29, 2026
- @needmorevram — Sep 29, 2026
- @0xSero — Sep 29, 2026
- @dangerm00se — Sep 29, 2026
- @0xSero — Sep 29, 2026
- @superalesha — Sep 29, 2026
- @0xSero — Sep 29, 2026
- @MiaAI_lab — Sep 29, 2026
- @jayleaton — Sep 29, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.