X local-model bench digest — Oct 7, 2026
As of Oct 7, 2026 — 27 posts collected from X. 59 measured runs catalogued, 2 published to the board, 25 kept to the digest only. Latest source post Oct 7, 2026 01:58 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.
DGX Spark × Qwen3.8-Flash-Next
Oct 7, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | — | 37.0 tok/s | — |
| — | NVFP4 | — | 21.5 tok/s | — |
| 400K | NVFP4 | — | — | 1495.0 tok/s |
| 32K | NVFP4 | — | — | 1769.0 tok/s |
Single DGX Spark owners have a new model worth paying attention to. Qwen3.8-Flash-Next NVFP4 can now be served on one Spark with a setup that is surprisingly capable for a single 128GB unified-memory machine. The headline numbers: → Up to 1M context → 1,431,164-token FP8 KV cache → Full image and video support → ~37 tok/s single-stream prose → ~86 tok/s across 4 concurrent streams → ~1,500 to 2,000 tok/s prefill →…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
TPU v5e-8 × Qwen3.8-27B
Oct 7, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262144 | BF16 | — | 130.0 tok/s | — |
| 262144 | BF16 | — | 78.0 tok/s | — |
| 262144 | BF16 | — | — | 10300.0 tok/s |
| 262144 | BF16 | — | 540.0 tok/s | — |
Qwen3.8-27B is running in full BF16 at around 130 tok/s on a free Kaggle TPU. No Q2. No Q4. No aggressive quantization. No expensive GPU rental. The setup uses a TPU v5e-8 and runs the full BF16 model across the TPU, with its native 262K context. The published measurements are roughly: ~130 tok/s decode with MTP ~78 tok/s without MTP ~10,300 tok/s prefill 262K context ~540 tok/s aggregate across 8 streams And that…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
Intel Arc Pro B70 × Qwen3.8-Flash-Next
Oct 6, 2026 · by @0xSero
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 24.0 tok/s | 1056.0 tok/s |
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 28.7 tok/s | 1092.0 tok/s |
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 26.4 tok/s | 1233.0 tok/s |
| 32768 | EXL3 3.05 bpw | sglang 24c872759256 | 23.9 tok/s | 1106.0 tok/s |
| 32768 | EXL3 3.05 bpw | sglang 24c872759256 | 25.4 tok/s | 1260.0 tok/s |
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 34.5 tok/s | 1018.0 tok/s |
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 54.8 tok/s | 1013.0 tok/s |
| 8192 | EXL3 3.05 bpw | sglang 24c872759256 | 76.3 tok/s | 1123.0 tok/s |
| 32768 | EXL3 3.05 bpw | sglang 24c872759256 | 34.1 tok/s | 1120.0 tok/s |
| 32768 | EXL3 3.05 bpw | sglang 24c872759256 | 42.9 tok/s | 1136.0 tok/s |
Qwen3.8-next-flash As promised, 1 GPU, 32GB of DDR4 and 50GB of NVMe. Getting: - 24 tok/s decode - 1106 tok/s prefill - 50 tok/s c4 https://github.com/0xSero/qwen38-flash-next-b70-offload/ https://t.co/9VmlprfT07
Full methodology stated (engine, quantization, context, hardware state, method) — 2 added to the benchmark board (2 of 10 measured rows queued; the rest are published here only).
RTX PRO 6000 x8 × GLM-5.3-NVFP4
Oct 6, 2026 · by @AiXsatoshi
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | — | 47.6 tok/s | — |
GLM-5.3-NVFP4 を2台のPCつなげて動かした パラメータ約753B/40B active、サイズ464GB Single:47.65 tok/s - 4 GPU×2ノードのTP4×PP2 - RTX PRO 6000 -pl 250W - ノード間は10GbE - MTPなし - overlap scheduleなし - radix cacheなし https://t.co/VruyPAV25q
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, measurement method).
Tesla V100 32GB × GLM 5.3 Flash
Oct 6, 2026 · by @PeasantSmith
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ1_S | unslothai | 25.0 tok/s | — |
GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM -18B Active parameters -25 tok/s peak decode -@UnslothAI IQ1_S (93GB file) -CPU I5-12600T +DDR5 5200 + Gen4 NVME -Custom strata The Volta GPU is probably the floor, results should be better with more RAM or newer GPUs https://t.co/4u37lnhGc2
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
Unspecified platform × Mistral Large
Oct 6, 2026 · by @miniroutersh · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 524K | — | — | 116.0 tok/s | — |
Mistral Large 4 (Le Chonk) is now live on Minirouter. → 1 trillion-parameters → 524k token context window → Open weights → "Made in Europe" → Quite fast (116 tok/s) → #64 in intelligence ranking
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
Unspecified platform × Deepseek V4.1 Flash
Oct 6, 2026 · by @davideciffa · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 50.0 tok/s | — |
Now we support Deepseek V4.1 Flash! Almost 50 tok/s for coding with speculative decoding. R9700 makes the Gorgon up to 3x faster 🏎️
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Unspecified platform × DeepSeek V4.1 Flash
Oct 6, 2026 · by @pupposandro · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 49.0 tok/s | — |
One of the main pieces of feedback we got on the Lucebox Zero 495 is that it's not clear why it matters to have both unified memory and a discrete GPU in the same machine. It matters a lot. In the article below, we show how the R9700 gets you 3x the speed on a big model like DeepSeek V4.1 Flash. The Lucebox Zero 495 runs the model across three tiers of memory: the 32 GB on the Radeon AI PRO R9700, the unified memory…
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
L40 x8 × GLM-5.3-Flash
Oct 6, 2026 · by @shadeformai
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 230K | — | — | 800.0 tok/s | — |
Officially, GLM-5.3-Flash requires Hopper GPUs or newer. However, on one 8x L40 node with no GPU peer-to-peer, we were able to run eight ~230k token coding sessions at once, at up to ~800 tok/s combined. To get GLM-5.3-Flash to run on Ada generation GPUs, we replaced the model's Hopper-only sparse attention with Triton kernels. Our first working build was less successful, managing 7.2 tok/s with only room for a…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
DGX Spark
Oct 6, 2026 · by @landontgreen · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | — | 2000.0 tok/s |
Abliterated GLM 5.3 Flash - 2x DGX Sparks With Abliterated flag on there is a slight speed loss, but nothing crazy. Improvement in prefill speed as well, 2000 tok/s now :) https://t.co/U6o4pBgund
Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
PCI Gen3 x8 + Imagine PCI Gen4 x16 × Qwen Flash Next
Oct 6, 2026 · by @needmorevram
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | Q4_K_XL | unsloth | — | 5000.0 tok/s |
Testing multi gpu setup with Strata 🤯 ~5000 Tok/s Prompt processing. Unsloth Qwen Flash Next Q4_K_XL Running on PCI Gen3 x8, Imagine PCI Gen4 x16 https://t.co/j3XBVsTEzI
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
Unspecified platform × Qwen3.8-Flash-Next
Oct 6, 2026 · by @EugeneSmarts · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 53.5 tok/s | — |
An AI agent was asked to check if an open-source bug already had an issue. Instead of checking, it went ahead and created one. Watching Qwen3.8-Flash-Next run tool calls like this highlights the fundamental tension in local agent design. Pushing local decode from 17.2 to 53.5 tok/s on cached turns makes long agent sessions feel snappy on a modest 16GB setup. But when an agent has zero impulse control, higher speed…
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
Also in this window
The 12 most recent posts are shown in full above; 15 posts from the same window are listed here in short form, each with its source post.
- M1 Max × Qwen3.8-Flash-Next-2bpw — 28.0 tok/s decode (1 measurement) · @Beamsters1
- GPU (32GB) × Qwen3.8-Next-Flash — 20.0 tok/s decode (1 measurement) · @NFT_Chen
- M5 Ultra Mac Studio x2 × Qwen 3.8 Flash Next — 271.0 tok/s decode (2 measurements) · @MiaAI_lab
- Tesla V100 × GLM-5.3-Flash — 44.0 tok/s decode (3 measurements) · @PeasantSmith
- DGX Spark x2 × Qwen3.8-Flash-Next — 1008.0 tok/s decode (4 measurements) · @vr8vr8
- RTX 3090 x20 × Qwen 397B — 20.0 tok/s decode (3 measurements) · @mylifcc
- DGX Spark x2 × GLM-5.3-Flash — 42.5 tok/s decode (3 measurements) · @sudoingX
- M5 Ultra Mac Studio × Qwen 3.8 27B — 21.4 tok/s decode (1 measurement) · @ashen_one
- RTX 3060 × qwen — 50.0 tok/s decode (1 measurement) · @sudoingX
- M5 Max MacBook Pro × Qwen3.8-27B — 70.0 tok/s decode (1 measurement) · @Oluwaphilemon1
- DGX Spark x2 × GLM-5.3-Flash — 40.0 tok/s decode (2 measurements) · @jurlycat
- DGX Spark × RED-SNOW-5.3-FLASH — 61.7 tok/s decode (3 measurements) · @ViC305
- DGX Spark × GLM-5.3-Flash — 32.0 tok/s decode (3 measurements) · @NFT_Chen
- Unspecified platform × TP5 GLM 5.3 — 272.0 tok/s decode (3 measurements) · @BrandonMusicKy
- Unspecified platform — 30.0 tok/s decode (1 measurement) · @cholf5
How to read this page
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @Oluwaphilemon1 — Oct 7, 2026
- @Oluwaphilemon1 — Oct 7, 2026
- @0xSero — Oct 6, 2026
- @AiXsatoshi — Oct 6, 2026
- @PeasantSmith — Oct 6, 2026
- @miniroutersh — Oct 6, 2026
- @davideciffa — Oct 6, 2026
- @pupposandro — Oct 6, 2026
- @shadeformai — Oct 6, 2026
- @landontgreen — Oct 6, 2026
- @needmorevram — Oct 6, 2026
- @EugeneSmarts — Oct 6, 2026
- @Beamsters1 — Oct 6, 2026
- @NFT_Chen — Oct 6, 2026
- @MiaAI_lab — Oct 6, 2026
- @PeasantSmith — Oct 6, 2026
- @vr8vr8 — Oct 6, 2026
- @mylifcc — Oct 6, 2026
- @sudoingX — Oct 6, 2026
- @ashen_one — Oct 6, 2026
- @sudoingX — Oct 6, 2026
- @Oluwaphilemon1 — Oct 6, 2026
- @jurlycat — Oct 6, 2026
- @ViC305 — Oct 6, 2026
- @NFT_Chen — Oct 6, 2026
- @BrandonMusicKy — Oct 6, 2026
- @cholf5 — Oct 6, 2026
FAQ
Are the numbers in this digest sourced?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a row to reach the sourced tier?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.