X local-model bench digest — Sep 28, 2026
As of Sep 28, 2026 — 12 posts collected from X. 41 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 28, 2026 07:34 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 5 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
DGX Spark x2 × Qwen3.8-Flash-Next-hibrid48
Sep 28, 2026 · by @vr8vr8
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | vllm 0.30 | 99.0 tok/s | — |
| — | NVFP4 | vllm 0.30 | 793.0 tok/s | — |
| 1K | NVFP4 | vllm 0.30 | 99.0 tok/s | 2543.0 tok/s |
| 128K | NVFP4 | vllm 0.30 | 89.8 tok/s | 3233.0 tok/s |
| 256K | NVFP4 | vllm 0.30 | 81.6 tok/s | 3094.0 tok/s |
| — | NVFP4 | vllm 0.30 | 80.0 tok/s | — |
Qwen3.8-Flash-Next - 4.1v🥳 🚀 Peak c=1 @ 118 tok/s 📊 Average c=1 @ 80 tok/s 🏭 Peak c=64 @ 834 tok/s ⚙️ Prefill 3,233 tok/s 💻 TTFT 0.42s ⚡️ https://github.com/myllmbox/qwen38-flash-next-cluster-recipe https://t.co/JpR5pYH4Ee
Full methodology stated (engine, quantization, context, hardware state, method) — 3 added to the benchmark board (3 of 6 measured rows queued; the rest are published here only).
M5 Ultra 80 core GPU × Qwen 3.8 Flash Next
Sep 28, 2026 · by @iam_agg
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 32768 | Q4 | tensorfold 0.3.4.1 | 152.0 tok/s | 2086.0 tok/s |
| 32768 | Q4 | mlx 26.9.6 | 124.0 tok/s | 3863.0 tok/s |
| 32768 | Q4 | mlx 0.7.0rc1 | 105.0 tok/s | 3694.0 tok/s |
Rough M5 Ultra 80 core GPU 256GB benches: Qwen 3.8 Flash Next, Q4 variants, MTP, 1 stream, 32k context, averaged code/prose: TensorFold 0.3.4.1: 152 tok/s decode, 2086 tok/s prefill mlx-serve 26.9.6: 124 tok/s, 3863 tok/s prefill oMLX 0.7.0rc1: 105 tok/s, 3694 tok/s prefill https://t.co/ZhSKeUYxWa
Digest-only: incomplete methodology (hardware state).
RTX 3060 × Qwen3.8-35B-A3B
Sep 28, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262144 | Q4_K_M | llama.cpp | — | 550.0 tok/s |
| 262144 | Q4_K_M | llama.cpp | 50.0 tok/s | — |
Qwen3.8-35B-A3B running at around 50 tok/s on an RTX 3060 12GB is the part of local AI that I find genuinely interesting. Not because 50 tok/s is some ridiculous inference record. Because this is a 35B model running on a GPU that only has 12GB of VRAM. The setup is surprisingly ordinary: RTX 3060 12GB 16GB system RAM Q4_K_M 262K context ~550 tok/s prefill ~50 tok/s decode The model is Empero’s Qwen3.8-35B-A3B, a…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
M4 Max × Qwen3.8-27B
Sep 28, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | tensorfold 0.3.4 | 118.3 tok/s | — |
| — | — | tensorfold 0.3.0 | 28.5 tok/s | — |
| — | — | mlx MTP3 | 65.8 tok/s | — |
| 31K | — | tensorfold 0.3.4 | 69.0 tok/s | — |
TensorFold 0.3.4 just shipped with M1–M4 support, and the M4 Max numbers are kind of ridiculous. Earlier this morning, Qwen3.8-27B on the same Mac Studio M4 Max was doing 28.5 tok/s with TensorFold 0.3.0. Tonight, TensorFold 0.3.4 is hitting: 118.3 tok/s Same Qwen3.8-27B. Same M4 Max. Same three prompts. Thinking off. Median of five runs. That’s a 4.0x jump from the drafting path. The serial floor is around 29.3…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization).
RTX 5070 x1 × Qwen3.8-Flash-Next
Sep 28, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 128K | Q2_0 | strata | 65.1 tok/s | — |
| 128K | Q2_0 | strata | — | 543.0 tok/s |
| 128K | IQ2_XS | strata | 52.0 tok/s | — |
| 128K | IQ2_XS | strata | — | 472.0 tok/s |
| 128K | IQ3_XXS | strata | 44.8 tok/s | — |
| 128K | IQ3_XXS | strata | — | 414.0 tok/s |
| — | — | llama.cpp | 15.0 tok/s | — |
Qwen3.8-Flash-Next was running at around 15 tok/s on a 12GB RTX 5070. Apparently, that wasn’t good enough. So instead of swapping the GPU or giving up on the model, he built his own inference engine around it. That’s how Strata came about. And the reported jump is pretty ridiculous: llama.cpp: ~15 tok/s Strata: up to 65.1 tok/s That’s more than 4x the decode speed on the same machine. The hardware isn’t some massive…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 5060 Ti × Qwen3.8-27B
Sep 28, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | — | llama.cpp v1.1 | 67.0 tok/s | — |
| 39K | — | llama.cpp v1.1 | 39.5 tok/s | — |
| 119K | — | llama.cpp v1.1 | 21.0 tok/s | — |
| 262K | — | llama.cpp v1.1 | 53.0 tok/s | — |
| 262K | — | llama.cpp v1.1 | 42.0 tok/s | — |
| 262K | — | llama.cpp v1.1 | 71.0 tok/s | — |
| 262K | — | llama.cpp v1.1 | 66.0 tok/s | — |
| 262K | — | llama.cpp v1.1 | 127.0 tok/s | — |
| 119K | — | llama.cpp v1.1 | — | 1060.0 tok/s |
Qwen3.8-27B is now doing something wild on a single RTX 5060 Ti 16GB. Bonsai 2, PrismML’s roughly 1.75-bit compression of Qwen3.8-27B, can run with: 262K context MTP head Vision All loaded simultaneously And the whole setup fits in about 15.2GB of the card’s 15.9GB usable VRAM. The reported speed is around 67 tok/s with the MTP head. For comparison: 67 tok/s with MTP 71 tok/s for code 53 tok/s without MTP 42 tok/s…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (measurement method).
DGX Spark × Qwen3.8-27B
Sep 28, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | custom inference engine | 45.0 tok/s | — |
Qwen3.8-27B NVFP4 is now hitting 45 tok/s on a single DGX Spark, running on a custom inference engine. And the part I care about most isn’t just the speed. It’s that the output is lossless. Every token produced by the accelerated path matches what plain decoding would have produced. So this isn’t a case of making the model look faster by changing the output or using a different decoding path. The model is still…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark x3 × GLM-5.3 Flash
Sep 28, 2026 · by @jakeharrisdev
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | jspark3 v1.8 | 73.3 tok/s | — |
| — | — | jspark3 v1.8 | 141.7 tok/s | — |
| — | — | jspark3 v1.8 | — | 1262.6 tok/s |
JSpark3 v1.8. GLM-5.3 Flash on three DGX Sparks, stock weights. 73.3 tok/s code on a single stream. 141.7 across 4 streams. 1,262.6 tok/s prefill. Everything's open: https://jakejh.com/jspark3/glm/?v=1.8 https://t.co/cdrrX73EvU
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).
Unspecified platform × GLM-5.3 Flash
Sep 27, 2026 · by @jakeharrisdev
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | jspark3 v1.8 | 73.3 tok/s | — |
| — | — | jspark3 v1.8 | 49.1 tok/s | — |
GLM-5.3 Flash on JSpark3 v1.8 hits 73.3 tok/s on code and 49.1 tok/s on prose. Stock weights, pinned recipe, every number measured. Recipe and card on Hugging Face: https://huggingface.co/jakejharris/jspark3
Digest-only: hardware.name missing; incomplete methodology (context length, hardware state).
DGX Spark x3 × GLM-5.3 Flash
Sep 27, 2026 · by @jakeharrisdev
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | jspark3 v1.8 | 141.7 tok/s | — |
| — | — | jspark3 v1.8 | 90.7 tok/s | — |
JSpark3 v1.8 is live. A big jump from v1.1, built with my agent fleet over the past few weeks. GLM-5.3 Flash. Up to 141.7 tok/s code and 90.7 tok/s prose decode at 4 streams on three DGX Sparks, stock weights. Pinned recipe, measured results. https://www.jakejh.com/jspark3/glm/?v=1.8
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state).
GB10 × GLM-5.3-Flash-EXL3-2.25bpw-sm121
Sep 27, 2026 · by @mr_r0b0t
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 8192 | EXL3 2.25bpw + DFlash2 3.00bpw | exllamav2 native | 50.0 tok/s | 22.6 tok/s |
r0b0tlab/GLM-5.3-Flash-EXL3-2.25bpw-sm121 + 3.00bpw DFlash2 drafter + native ExLlamaV3 runtime! Tested on a single @NVIDIAAI GB10 ♥️ At 8K C1, code: 22.61→49.96 tok/s with K=5 Exact single-key retrieval at 259,993 prompt tokens Extended eval suite in progress 🤓 https://t.co/oYxwpC4I7S
Digest-only: incomplete methodology (engine+version, hardware state, measurement method).
DGX Spark × Qwen3.8-27B
Sep 27, 2026 · by @0xBakeer
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | my own engine | 45.0 tok/s | — |
Qwen3.8-27B NVFP4 on one DGX Spark, running on my own engine. 45 tok/s on prose, and it's lossless: every token is exactly what plain decoding would write. All the other numbers are live in this 2 min demo. Still a WIP 🚧 I'll publish it this week. Really proud of this one https://t.co/OGRFYhoh35
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @vr8vr8 — Sep 28, 2026
- @iam_agg — Sep 28, 2026
- @Oluwaphilemon1 — Sep 28, 2026
- @Oluwaphilemon1 — Sep 28, 2026
- @Oluwaphilemon1 — Sep 28, 2026
- @Oluwaphilemon1 — Sep 28, 2026
- @Oluwaphilemon1 — Sep 28, 2026
- @jakeharrisdev — Sep 28, 2026
- @jakeharrisdev — Sep 27, 2026
- @jakeharrisdev — Sep 27, 2026
- @mr_r0b0t — Sep 27, 2026
- @0xBakeer — Sep 27, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.