DiggTop ⌕

X local-model bench digest — Oct 7, 2026

As of Oct 7, 2026 — 27 posts collected from X. 59 measured runs catalogued, 2 published to the board, 25 kept to the digest only. Latest source post Oct 7, 2026 01:58 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.

DGX Spark × Qwen3.8-Flash-Next

Oct 7, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— NVFP4 — 37.0 tok/s —
— NVFP4 — 21.5 tok/s —
400K NVFP4 — — 1495.0 tok/s
32K NVFP4 — — 1769.0 tok/s

Single DGX Spark owners have a new model worth paying attention to. Qwen3.8-Flash-Next NVFP4 can now be served on one Spark with a setup that is surprisingly capable for a single 128GB unified-memory machine. The headline numbers: → Up to 1M context → 1,431,164-token FP8 KV cache → Full image and video support → ~37 tok/s single-stream prose → ~86 tok/s across 4 concurrent streams → ~1,500 to 2,000 tok/s prefill →…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

TPU v5e-8 × Qwen3.8-27B

Oct 7, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 BF16 — 130.0 tok/s —
262144 BF16 — 78.0 tok/s —
262144 BF16 — — 10300.0 tok/s
262144 BF16 — 540.0 tok/s —

Qwen3.8-27B is running in full BF16 at around 130 tok/s on a free Kaggle TPU. No Q2. No Q4. No aggressive quantization. No expensive GPU rental. The setup uses a TPU v5e-8 and runs the full BF16 model across the TPU, with its native 262K context. The published measurements are roughly: ~130 tok/s decode with MTP ~78 tok/s without MTP ~10,300 tok/s prefill 262K context ~540 tok/s aggregate across 8 streams And that…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Intel Arc Pro B70 × Qwen3.8-Flash-Next

Oct 6, 2026 · by @0xSero

Context Quant Engine Decode Prefill
8192 EXL3 3.05 bpw sglang 24c872759256 24.0 tok/s 1056.0 tok/s
8192 EXL3 3.05 bpw sglang 24c872759256 28.7 tok/s 1092.0 tok/s
8192 EXL3 3.05 bpw sglang 24c872759256 26.4 tok/s 1233.0 tok/s
32768 EXL3 3.05 bpw sglang 24c872759256 23.9 tok/s 1106.0 tok/s
32768 EXL3 3.05 bpw sglang 24c872759256 25.4 tok/s 1260.0 tok/s
8192 EXL3 3.05 bpw sglang 24c872759256 34.5 tok/s 1018.0 tok/s
8192 EXL3 3.05 bpw sglang 24c872759256 54.8 tok/s 1013.0 tok/s
8192 EXL3 3.05 bpw sglang 24c872759256 76.3 tok/s 1123.0 tok/s
32768 EXL3 3.05 bpw sglang 24c872759256 34.1 tok/s 1120.0 tok/s
32768 EXL3 3.05 bpw sglang 24c872759256 42.9 tok/s 1136.0 tok/s

Qwen3.8-next-flash As promised, 1 GPU, 32GB of DDR4 and 50GB of NVMe. Getting: - 24 tok/s decode - 1106 tok/s prefill - 50 tok/s c4 https://github.com/0xSero/qwen38-flash-next-b70-offload/ https://t.co/9VmlprfT07

Full methodology stated (engine, quantization, context, hardware state, method) — 2 added to the benchmark board (2 of 10 measured rows queued; the rest are published here only).

RTX PRO 6000 x8 × GLM-5.3-NVFP4

Oct 6, 2026 · by @AiXsatoshi

Context Quant Engine Decode Prefill
— NVFP4 — 47.6 tok/s —

GLM-5.3-NVFP4 を2台のPCつなげて動かした パラメータ約753B/40B active、サイズ464GB Single:47.65 tok/s - 4 GPU×2ノードのTP4×PP2 - RTX PRO 6000 -pl 250W - ノード間は10GbE - MTPなし - overlap scheduleなし - radix cacheなし https://t.co/VruyPAV25q

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, measurement method).

Tesla V100 32GB × GLM 5.3 Flash

Oct 6, 2026 · by @PeasantSmith

Context Quant Engine Decode Prefill
— IQ1_S unslothai 25.0 tok/s —

GLM 5.3 Flash on a single Tesla V100 32GB + 64GB of RAM -18B Active parameters -25 tok/s peak decode -@UnslothAI IQ1_S (93GB file) -CPU I5-12600T +DDR5 5200 + Gen4 NVME -Custom strata The Volta GPU is probably the floor, results should be better with more RAM or newer GPUs https://t.co/4u37lnhGc2

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform × Mistral Large

Oct 6, 2026 · by @miniroutersh · third-party report

Context Quant Engine Decode Prefill
524K — — 116.0 tok/s —

Mistral Large 4 (Le Chonk) is now live on Minirouter. → 1 trillion-parameters → 524k token context window → Open weights → "Made in Europe" → Quite fast (116 tok/s) → #64 in intelligence ranking

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

Unspecified platform × Deepseek V4.1 Flash

Oct 6, 2026 · by @davideciffa · third-party report

Context Quant Engine Decode Prefill
— — — 50.0 tok/s —

Now we support Deepseek V4.1 Flash! Almost 50 tok/s for coding with speculative decoding. R9700 makes the Gorgon up to 3x faster 🏎️

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Unspecified platform × DeepSeek V4.1 Flash

Oct 6, 2026 · by @pupposandro · third-party report

Context Quant Engine Decode Prefill
— — — 49.0 tok/s —

One of the main pieces of feedback we got on the Lucebox Zero 495 is that it's not clear why it matters to have both unified memory and a discrete GPU in the same machine. It matters a lot. In the article below, we show how the R9700 gets you 3x the speed on a big model like DeepSeek V4.1 Flash. The Lucebox Zero 495 runs the model across three tiers of memory: the 32 GB on the Radeon AI PRO R9700, the unified memory…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

L40 x8 × GLM-5.3-Flash

Oct 6, 2026 · by @shadeformai

Context Quant Engine Decode Prefill
230K — — 800.0 tok/s —

Officially, GLM-5.3-Flash requires Hopper GPUs or newer. However, on one 8x L40 node with no GPU peer-to-peer, we were able to run eight ~230k token coding sessions at once, at up to ~800 tok/s combined. To get GLM-5.3-Flash to run on Ada generation GPUs, we replaced the model's Hopper-only sparse attention with Triton kernels. Our first working build was less successful, managing 7.2 tok/s with only room for a…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).

DGX Spark

Oct 6, 2026 · by @landontgreen · third-party report

Context Quant Engine Decode Prefill
— — — — 2000.0 tok/s

Abliterated GLM 5.3 Flash - 2x DGX Sparks With Abliterated flag on there is a slight speed loss, but nothing crazy. Improvement in prefill speed as well, 2000 tok/s now :) https://t.co/U6o4pBgund

Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

PCI Gen3 x8 + Imagine PCI Gen4 x16 × Qwen Flash Next

Oct 6, 2026 · by @needmorevram

Context Quant Engine Decode Prefill
— Q4_K_XL unsloth — 5000.0 tok/s

Testing multi gpu setup with Strata 🤯 ~5000 Tok/s Prompt processing. Unsloth Qwen Flash Next Q4_K_XL Running on PCI Gen3 x8, Imagine PCI Gen4 x16 https://t.co/j3XBVsTEzI

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Oct 6, 2026 · by @EugeneSmarts · third-party report

Context Quant Engine Decode Prefill
— — — 53.5 tok/s —

An AI agent was asked to check if an open-source bug already had an issue. Instead of checking, it went ahead and created one. Watching Qwen3.8-Flash-Next run tool calls like this highlights the fundamental tension in local agent design. Pushing local decode from 17.2 to 53.5 tok/s on cached turns makes long agent sessions feel snappy on a modest 16GB setup. But when an agent has zero impulse control, higher speed…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Also in this window

The 12 most recent posts are shown in full above; 15 posts from the same window are listed here in short form, each with its source post.

How to read this page

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest sourced?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a row to reach the sourced tier?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑