X local-model bench digest — Oct 2, 2026
As of Oct 2, 2026 — 12 posts collected from X. 32 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 2, 2026 11:05 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 35 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
RTX 3090 x1 × Deepseek-v4.1-flash
Oct 2, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 18.0 tok/s | — |
| — | — | — | — | 700.0 tok/s |
Deepseek-v4.1-flash on 1x 3090 + 256GB of DDR4 + 200GB of NVMe - 15 -> 21 tok/s decode - 600 -> 800 tok/s prefill down to 6000$ https://t.co/IPXBmJ7XtH
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x4 × GLM-5.3
Oct 2, 2026 · by @majewskizby
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 32K | Int4/Int8 mixed | sparkdash | 30.0 tok/s | — |
| 32K | Int4/Int8 mixed | rigmark 1.0.0 | 24.9 tok/s | — |
| 32K | Int4/Int8 mixed | sparkdash | 61.1 tok/s | — |
| 4K | Int4/Int8 mixed | — | — | 1000.0 tok/s |
| 32K | Int4/Int8 mixed | — | — | 840.0 tok/s |
Full GLM-5.3 (753B) TP4 on 4× DGX Spark: initial recipe 🐓 • Prose decode: 30.0 tok/s at c1 (sparkDash), 24.9 tok/s with thinking on (RigMark) • Prose aggregate: 61.1 tok/s at c4 • Prefill: ~1.0k tok/s at 4k, 840 at 32k • 32k context, ~44k-token FP8 KV cache • Int4/Int8 mixed weights, native MTP (K=2) speculative decoding • Quality gate: 72.7/75 sparkDash: thinking off, 256 output tokens. RigMark 1.0.0: thinking on…
Digest-only: incomplete methodology (hardware state, measurement method).
DGX Spark x3 × GLM 5.3 Flash
Oct 2, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | tensorfold | 77.0 tok/s | — |
| — | EXL3 | tensorfold | 146.0 tok/s | — |
GLM 5.3 Flash EXL3 TensorFold on 3x DGX Sparks - 77 tok/s on prose, single stream - 146 tok/s on prose, 4 streams - 6M KV cache Coming soon. https://t.co/Q66IxrGxh9
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
M5 Ultra × Qwen3.8-Flash-Next
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | mlx 26.10.1 | — | 5400.0 tok/s |
Qwen3.8-27B just got significantly faster on Apple Silicon, and the model itself didn’t change. MLX-Serve 26.10.1 is out with major inference improvements across the board. With the Qwen3.8-27B drafter enabled: M5 Ultra: +66% M4 Max: +28% M1 Pro: +37% That’s a huge performance jump from a runtime update alone. And Qwen3.8-Flash-Next gets an impressive boost too: M4 Max: +11% decode M5 Ultra: +51% prefill The M5…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).
custom build × DeeepSeek-v4.1-Flash
Oct 2, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 21.0 tok/s | — |
| — | — | — | — | 100.0 tok/s |
Rejoice! DeeepSeek-v4.1-Flash running on 200GB of NVMe + 400GB of DDR4 + 24GB of VRAM Caps out at 21 tok/s decode, prefill is 100 tok/s but working on that now Will have a working image by EOD In total this deepseek running on 9000$ of hardware, 2 sparks recipe also coming https://t.co/e52NWlAUkg
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Unspecified platform
Oct 2, 2026 · by @JoelDeTeves · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | AWQ | — | 180.0 tok/s | — |
Would you believe @ornith 1.5 35B did this in a single shot? The little 35B model that could - ripping upwards of 180 tok/s with DFLASH2 AWQ quant by @cyankiwi_ai - running on the RTX A6000 (Ampere) 48 GB https://huggingface.co/cyankiwi/Ornith-1.5-35B-A3B-AWQ-INT4 DFLASH drafter available here https://t.co/rwgXWSvGA9
Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).
M5 Ultra × Qwen3.8 Flash Next
Oct 2, 2026 · by @plotarmordev
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | 4bit | tensorfold 0.6.0 | 199.0 tok/s | — |
| — | 4bit | tensorfold 0.6.0 | 88.0 tok/s | — |
| — | 4bit | tensorfold 0.6.0 | 358.0 tok/s | — |
| — | 4bit | tensorfold 0.6.0 | 255.0 tok/s | — |
The M5 Ultra has 4.4x the memory bandwidth of a DGX Spark. Under load, it's only 1.4x faster! Same engine (TensorFold 0.6.0), same 4bit Qwen3.8 Flash Next, same benchmark, one Spark vs one Mac: - 1 session: Mac 2.3x faster (199 vs 88 tok/s) - 8 sessions: 1.4x (358 vs 255 total) - Reading long prompts: ~1.2x The Mac wins at one session. BUT under load and on long prompts, the Spark gets much closer than the spec…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).
RTX 5090 × Qwen3.8-Flash-Next
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | strata | 135.0 tok/s | — |
| — | — | strata | — | 4000.0 tok/s |
Is it a little late to say this? Strata is seriously impressive. Qwen3.8-Flash-Next is an enormous model. The quantized model file can be around 87GB, which sounds like the kind of thing that should immediately rule out a normal desktop. Except it doesn’t. With Strata, a single RTX 5090 can reportedly push roughly: 120–150 tok/s decode ~4,000 tok/s prefill And that’s with Qwen3.8-Flash-Next running locally. The…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Unspecified platform × Qwen3.8-Flash-Next
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1M | — | strata | 155.0 tok/s | — |
| 1M | — | strata | — | 4000.0 tok/s |
Strata + Qwen3.8-Flash-Next is starting to make 1 million tokens of context feel almost normal. And honestly, this is messing with my head. The reported setup is reaching roughly: 150–160 tok/s decode ~4,000 tok/s prefill ~1M token context And pushing the context from the usual range toward 1 million tokens only adds around 10GB of system memory. That last number is the part I keep coming back to. Because…
Digest-only: hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
RTX 3090 x1 × Qwen3.8-27B
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 25K | — | vllm | 381.0 tok/s | — |
| 25K | — | vllm | 260.0 tok/s | — |
| — | — | vllm | 133.0 tok/s | — |
A single RTX 3090 just pushed Qwen3.8-27B to around 381 tok/s. 24GB of VRAM. And no, this isn’t a normal-chat benchmark. The crazy number comes from changing how speculative decoding works. The progression is already impressive: ~82 tok/s: initial single-user setup ~114 tok/s: optimized MTP ~138 tok/s: DFlash2 + lookup drafting ~381 tok/s: longer verification blocks + context lookup The hardware stayed the same: 1x…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
RTX 4090 × Qwen3.8-27B
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 30K | Q4_K_M | llama.cpp | — | 1725.0 tok/s |
| 30K | Q4_K_M | llama.cpp | 87.0 tok/s | — |
| 80K | Q4_K_M | llama.cpp | — | 1789.0 tok/s |
| 80K | Q4_K_M | llama.cpp | 84.2 tok/s | — |
| 110K | Q4_K_M | llama.cpp | — | 1767.0 tok/s |
| 110K | Q4_K_M | llama.cpp | 83.3 tok/s | — |
Qwen3.8-27B just hit 90 tok/s on a single RTX 4090. 24GB of VRAM. And the interesting part isn’t a new model. It’s DFlash 2. The setup uses Qwen3.8-27B Q4_K_M with Z-Lab’s DFlash 2 drafter, paired with Unsloth’s UD-Q4_K_XL quant and a patched llama.cpp build. The previous setup was already doing around 60 tok/s with native MTP at roughly 130K context. DFlash 2 pushes that much further, with reported decode speeds…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
RTX 3060 × Bonsai 2 27B
Oct 2, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 125K | — | hermes agent | 22.0 tok/s | — |
| 125K | — | hermes agent | 50.0 tok/s | — |
This is what 12GB of VRAM can build in 2026. An RTX 3060 with 12GB of VRAM just spent five hours building an entire playable game locally. The setup: RTX 3060 12GB Bonsai 2 27B, based on Qwen3.8-27B 5.95GB weights Hermes Agent 125K context ~50 tok/s fresh ~22 tok/s average By the end of the session, the agent had generated: 328K tokens 8 JavaScript files 2,368 lines of code A complete playable game Zero handwritten…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @0xSero — Oct 2, 2026
- @majewskizby — Oct 2, 2026
- @MiaAI_lab — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
- @0xSero — Oct 2, 2026
- @JoelDeTeves — Oct 2, 2026
- @plotarmordev — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
- @Oluwaphilemon1 — Oct 2, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.