DiggTop ⌕

X local-model bench digest — Oct 2, 2026

As of Oct 2, 2026 — 12 posts collected from X. 32 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 2, 2026 11:05 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 35 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

RTX 3090 x1 × Deepseek-v4.1-flash

Oct 2, 2026 · by @0xSero

Context Quant Engine Decode Prefill
— — — 18.0 tok/s —
— — — — 700.0 tok/s

Deepseek-v4.1-flash on 1x 3090 + 256GB of DDR4 + 200GB of NVMe - 15 -> 21 tok/s decode - 600 -> 800 tok/s prefill down to 6000$ https://t.co/IPXBmJ7XtH

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x4 × GLM-5.3

Oct 2, 2026 · by @majewskizby

Context Quant Engine Decode Prefill
32K Int4/Int8 mixed sparkdash 30.0 tok/s —
32K Int4/Int8 mixed rigmark 1.0.0 24.9 tok/s —
32K Int4/Int8 mixed sparkdash 61.1 tok/s —
4K Int4/Int8 mixed — — 1000.0 tok/s
32K Int4/Int8 mixed — — 840.0 tok/s

Full GLM-5.3 (753B) TP4 on 4× DGX Spark: initial recipe 🐓 • Prose decode: 30.0 tok/s at c1 (sparkDash), 24.9 tok/s with thinking on (RigMark) • Prose aggregate: 61.1 tok/s at c4 • Prefill: ~1.0k tok/s at 4k, 840 at 32k • 32k context, ~44k-token FP8 KV cache • Int4/Int8 mixed weights, native MTP (K=2) speculative decoding • Quality gate: 72.7/75 sparkDash: thinking off, 256 output tokens. RigMark 1.0.0: thinking on…

Digest-only: incomplete methodology (hardware state, measurement method).

DGX Spark x3 × GLM 5.3 Flash

Oct 2, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
— EXL3 tensorfold 77.0 tok/s —
— EXL3 tensorfold 146.0 tok/s —

GLM 5.3 Flash EXL3 TensorFold on 3x DGX Sparks - 77 tok/s on prose, single stream - 146 tok/s on prose, 4 streams - 6M KV cache Coming soon. https://t.co/Q66IxrGxh9

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

M5 Ultra × Qwen3.8-Flash-Next

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — mlx 26.10.1 — 5400.0 tok/s

Qwen3.8-27B just got significantly faster on Apple Silicon, and the model itself didn’t change. MLX-Serve 26.10.1 is out with major inference improvements across the board. With the Qwen3.8-27B drafter enabled: M5 Ultra: +66% M4 Max: +28% M1 Pro: +37% That’s a huge performance jump from a runtime update alone. And Qwen3.8-Flash-Next gets an impressive boost too: M4 Max: +11% decode M5 Ultra: +51% prefill The M5…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).

custom build × DeeepSeek-v4.1-Flash

Oct 2, 2026 · by @0xSero

Context Quant Engine Decode Prefill
— — — 21.0 tok/s —
— — — — 100.0 tok/s

Rejoice! DeeepSeek-v4.1-Flash running on 200GB of NVMe + 400GB of DDR4 + 24GB of VRAM Caps out at 21 tok/s decode, prefill is 100 tok/s but working on that now Will have a working image by EOD In total this deepseek running on 9000$ of hardware, 2 sparks recipe also coming https://t.co/e52NWlAUkg

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform

Oct 2, 2026 · by @JoelDeTeves · third-party report

Context Quant Engine Decode Prefill
— AWQ — 180.0 tok/s —

Would you believe @ornith 1.5 35B did this in a single shot? The little 35B model that could - ripping upwards of 180 tok/s with DFLASH2 AWQ quant by @cyankiwi_ai - running on the RTX A6000 (Ampere) 48 GB https://huggingface.co/cyankiwi/Ornith-1.5-35B-A3B-AWQ-INT4 DFLASH drafter available here https://t.co/rwgXWSvGA9

Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

M5 Ultra × Qwen3.8 Flash Next

Oct 2, 2026 · by @plotarmordev

Context Quant Engine Decode Prefill
— 4bit tensorfold 0.6.0 199.0 tok/s —
— 4bit tensorfold 0.6.0 88.0 tok/s —
— 4bit tensorfold 0.6.0 358.0 tok/s —
— 4bit tensorfold 0.6.0 255.0 tok/s —

The M5 Ultra has 4.4x the memory bandwidth of a DGX Spark. Under load, it's only 1.4x faster! Same engine (TensorFold 0.6.0), same 4bit Qwen3.8 Flash Next, same benchmark, one Spark vs one Mac: - 1 session: Mac 2.3x faster (199 vs 88 tok/s) - 8 sessions: 1.4x (358 vs 255 total) - Reading long prompts: ~1.2x The Mac wins at one session. BUT under load and on long prompts, the Spark gets much closer than the spec…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).

RTX 5090 × Qwen3.8-Flash-Next

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — strata 135.0 tok/s —
— — strata — 4000.0 tok/s

Is it a little late to say this? Strata is seriously impressive. Qwen3.8-Flash-Next is an enormous model. The quantized model file can be around 87GB, which sounds like the kind of thing that should immediately rule out a normal desktop. Except it doesn’t. With Strata, a single RTX 5090 can reportedly push roughly: 120–150 tok/s decode ~4,000 tok/s prefill And that’s with Qwen3.8-Flash-Next running locally. The…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
1M — strata 155.0 tok/s —
1M — strata — 4000.0 tok/s

Strata + Qwen3.8-Flash-Next is starting to make 1 million tokens of context feel almost normal. And honestly, this is messing with my head. The reported setup is reaching roughly: 150–160 tok/s decode ~4,000 tok/s prefill ~1M token context And pushing the context from the usual range toward 1 million tokens only adds around 10GB of system memory. That last number is the part I keep coming back to. Because…

Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 3090 x1 × Qwen3.8-27B

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
25K — vllm 381.0 tok/s —
25K — vllm 260.0 tok/s —
— — vllm 133.0 tok/s —

A single RTX 3090 just pushed Qwen3.8-27B to around 381 tok/s. 24GB of VRAM. And no, this isn’t a normal-chat benchmark. The crazy number comes from changing how speculative decoding works. The progression is already impressive: ~82 tok/s: initial single-user setup ~114 tok/s: optimized MTP ~138 tok/s: DFlash2 + lookup drafting ~381 tok/s: longer verification blocks + context lookup The hardware stayed the same: 1x…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).

RTX 4090 × Qwen3.8-27B

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
30K Q4_K_M llama.cpp — 1725.0 tok/s
30K Q4_K_M llama.cpp 87.0 tok/s —
80K Q4_K_M llama.cpp — 1789.0 tok/s
80K Q4_K_M llama.cpp 84.2 tok/s —
110K Q4_K_M llama.cpp — 1767.0 tok/s
110K Q4_K_M llama.cpp 83.3 tok/s —

Qwen3.8-27B just hit 90 tok/s on a single RTX 4090. 24GB of VRAM. And the interesting part isn’t a new model. It’s DFlash 2. The setup uses Qwen3.8-27B Q4_K_M with Z-Lab’s DFlash 2 drafter, paired with Unsloth’s UD-Q4_K_XL quant and a patched llama.cpp build. The previous setup was already doing around 60 tok/s with native MTP at roughly 130K context. DFlash 2 pushes that much further, with reported decode speeds…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 3060 × Bonsai 2 27B

Oct 2, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
125K — hermes agent 22.0 tok/s —
125K — hermes agent 50.0 tok/s —

This is what 12GB of VRAM can build in 2026. An RTX 3060 with 12GB of VRAM just spent five hours building an entire playable game locally. The setup: RTX 3060 12GB Bonsai 2 27B, based on Qwen3.8-27B 5.95GB weights Hermes Agent 125K context ~50 tok/s fresh ~22 tok/s average By the end of the session, the agent had generated: 328K tokens 8 JavaScript files 2,368 lines of code A complete playable game Zero handwritten…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑