DiggTop ⌕

X local-model bench digest — Oct 1, 2026

As of Oct 1, 2026 — 12 posts collected from X. 38 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 1, 2026 12:01 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 16 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark × GLM-5.3

Oct 1, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
131K NVFP4 vllm 30.0 tok/s —

GLM-5.3 Flash just one shotted my entire floating tree benchmark on 2x DGX Spark, here is what it took and the hardware it ran on: 2x DGX Spark, NVIDIA's NVFP4 weights split over one ConnectX-7 cable, 131K context with vision on served with vLLM and MTP k=3, 20 to 30 tok/s decode hermes agent as the harness, one prompt 4+ hours, 8 files, 2,356 lines of three.js, zero lines written by hand all 12 tests green 1h34m…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × GLM 5.3 Flash

Oct 1, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
1M EXL3 tensorfold 60.0 tok/s —
1M EXL3 tensorfold 108.0 tok/s —
— EXL3 tensorfold — 1950.0 tok/s

Run GLM 5.3 Flash EXL3 with TensorFold ⚡️ This is a completely new recipe that ushers a whole new level of performance for @NVIDIAAI 2x DGX Sparks. Conservative default for stability: - 1M context by default - 2.7M KV cache pool (!) - Yes, it's not a typo - 2.7M KV in just two Sparks - 4 concurrent streams by default Performance: - 60 tok/s on prose, single stream. - 108 tok/s on prose, 4 concurrent streams. ~1,950…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × Qwen3.8-Flash-Next

Oct 1, 2026 · by @sfxnz

Context Quant Engine Decode Prefill
— NVFP4 vllm 0.30 57.0 tok/s —
— NVFP4 vllm 0.30 84.3 tok/s —
— NVFP4 vllm 0.30 156.0 tok/s —

Qwen3.8-Flash-Next NVFP4 on 2x DGX Spark, now on vLLM 0.30 ∙Single stream prose decode 36.9 → 57 tok/s (+54%) ∙Structued output 63.9 → 84.3 tok/s (+32%) ∙4.4x throughput at c=2 35.7 → 156 tok/s (fixed) ∙-86% turn-2 latency at 16k context (5.1s → 0.7s) Enjoy! https://t.co/eXrTHkj236

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, measurement method).

Unspecified platform × GLM-5.3

Oct 1, 2026 · by @NFT_Chen · third-party report

Context Quant Engine Decode Prefill
— FP8 — 670.0 tok/s —

🤨我靠!重写内核,魔改版 GLM-5.3 Flash 推理速度达到 670 tok/s! 主要亮点: 🔹670 tok/s(Vercel AI Gateway)
🔹$0.11 / 1M input,$0.45 / 1M output,$0.03 / 1M cached
🔹1M 上下文,原生 FP8
🔹已支持 AMD,同模型同 API
🔹近 24h 缓存命中率 99.7%
🔹兼容 OpenAI / Anthropic,支持图文、工具调用、JSON、流式
🔹零数据留存,不用于训练 上手 3 步: 1️⃣打开 https://t.co/CkqSFpEuoQ 注册并拿 API Key 2️⃣把现有 OpenAI SDK 的 baseURL 改成 RunInfra 端点,模型名保持 GLM-5.3 Flash 3️⃣也可走 Vercel AI Gateway:https://t.co/XIQ0fQpqyO 换个 URL…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

M3 Ultra × Qwen

Oct 1, 2026 · by @volatilemarkts · third-party report

Context Quant Engine Decode Prefill
— — mlx 5.3 60.0 tok/s —

TensorFold on my Mac Studio M3 Ultra, 5 days after launch: Qwen 27B code: +269% vs plain MLX GLM-5.3-Flash: +86-108% vs mlx-vlm Prompt reading: +170% My own GLM seat: 45 → 60 tok/s (+33%) Same answers, bit for bit. @Apple https://github.com/ashhart/TensorFold https://t.co/zV67faufli

Digest-only: confidence 0.35 below 0.60; incomplete methodology (quantization, context length, hardware state, measurement method).

RTX 5090 × Qwen3.8-Flash-Next

Oct 1, 2026 · by @Oluwaphilemon1 · third-party report

Context Quant Engine Decode Prefill
— — — 140.0 tok/s —
— — — 70.0 tok/s —

Qwen3.8-Flash-Next is already getting attention for its intelligence and local inference speed. With Strata on an RTX 5090, reported speeds can reach roughly 100-140 tok/s. Even an RTX 5070 Ti can reportedly push around 50-70 tok/s. So naturally, the next question is: What if you could run Qwen3.8-Flash-Next locally without the model’s built-in refusal behavior? I initially assumed that meant doing the usual work…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × Qwen3.8-Flash-Next

Oct 1, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
48800 — tensorfold 0.3.4 — 107.0 tok/s
48800 — tensorfold 0.3.6.3 — 1953.0 tok/s
48800 — tensorfold 0.3.4 97.4 tok/s —
48800 — tensorfold 0.3.6.3 108.5 tok/s —
48800 — tensorfold 0.3.4 96.9 tok/s —
48800 — tensorfold 0.3.6.3 107.7 tok/s —
48800 — tensorfold 0.3.4 113.1 tok/s —
48800 — tensorfold 0.3.6.3 129.1 tok/s —
48800 — tensorfold 0.3.4 81.7 tok/s —
48800 — tensorfold 0.3.6.3 87.9 tok/s —
48800 — tensorfold 0.3.4 36.1 tok/s —
48800 — tensorfold 0.3.6.3 39.3 tok/s —

Qwen3.8-Flash-Next just got a huge inference improvement on DGX Spark with TensorFold. Same model. Same hardware. Same workload. The difference is TensorFold 0.3.4 → 0.3.6.3. Setup: Qwen3.8-Flash-Next Vontra MLX 4-bit + MTP drafts TP1 131K context window Two DGX Sparks for the decode tests The biggest change is prefill. A 48.8K-token prompt used to take about 453 seconds on TensorFold 0.3.4. After the update: 453s →…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (hardware state).

DGX Spark × Qwen3.8-27B

Oct 1, 2026 · by @Oluwaphilemon1 · third-party report

Context Quant Engine Decode Prefill
150K — — 90.0 tok/s —

Qwen3.8-27B running locally on a single DGX Spark. And to be clear, this is the dense Qwen3.8-27B, not Qwen3.8-Flash. The setup is getting around 50-90 tok/s during decoding depending on the workload, while optimized caching algorithms and the StairCut method dramatically reduce the cost of working with a large context. The reported result is especially interesting around the ~150K context mark, where prefill…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Sep 30, 2026 · by @yume_arasaki

Context Quant Engine Decode Prefill
— EXL3 4-bit vllm 66.4 tok/s —
— EXL3 4-bit tensorfold 109.1 tok/s —
— EXL3 4-bit vllm 28.5 tok/s —
— EXL3 4-bit tensorfold 57.7 tok/s —
— EXL3 4-bit vllm 61.3 tok/s —
— EXL3 4-bit tensorfold 96.5 tok/s —
— EXL3 4-bit vllm 44.7 tok/s —
— EXL3 4-bit tensorfold 80.2 tok/s —

Last post I showed you how to get over 100 tok/s on two DGX Sparks running Qwen 3.8 Flash. Today I confirmed the a 2x speed-up is coming for Local AI. The answer is TensorFold. I'll tell you how I ran the smartest model on desk AI using this new invention. Same two Sparks. Same desk. Same 200 Gb/s cable between them. The model is GLM-5.3-Flash, an abliterated EXL3 build based on Orca Router's abliterated model, full…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state).

RTX 3090 x2 × Qwen3.8-Next-Flash

Sep 30, 2026 · by @RRunner144

Context Quant Engine Decode Prefill
— IQ3_S — — 1050.0 tok/s
— — — 100.0 tok/s —

+100 tok/s qwen 3.8 next flash 256k kv cache iq3_s context between 900-1200 tok/s 2x rtx 3090 and 80gb DDR4 salvaged HP Z840 w/old 12 core Xeons Spent all day yesterday optimizing a previous version of Strata. Sent Niko some notes on what I found to get dual 3090s online, he drops a newer revision same day and my full stack goes up +10-15%. My daily driver Hermes agent is 3X faster, smarter and runs cooler than my…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Consumer Intel Arc × Qwen3.8-27B

Sep 30, 2026 · by @0x0SojalSec

Context Quant Engine Decode Prefill
262K — — 60.0 tok/s —
262K — — — 1500.0 tok/s
262K — — — 6000.0 tok/s

Security researchers : should check this out, for bughunting. - 262k context, Vision, 60 tok/s, - 1,500–6,000 tok/s prefill. - Qwen3.8-27B on an Consumer Intel Arc. - Hunting bugs until the payout hits. https://t.co/WHgukRCBPH

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 3090 x1 × glm-5.3-flash

Sep 30, 2026 · by @0xSero

Context Quant Engine Decode Prefill
— exl3-3bpw exllamav2 23.4 tok/s —

23.4 tok/s glm-5.3-flash 131k kv cache exl3-3bpw context between 500-800 tok/s 1x rtx 3090 and 196gb DDR4 https://t.co/cLRE8vLqVV

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑