DiggTop ⌕

X local-model bench digest — Sep 30, 2026

As of Sep 30, 2026 — 12 posts collected from X. 35 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 30, 2026 11:34 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 20 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

RTX Pro 6000 x8 × GLM-5.3

Sep 30, 2026 · by @keennay

Context Quant Engine Decode Prefill
251K NVFP4 nvidia 54.0 tok/s —
123K NVFP4 incoai DFlash2 93.0 tok/s —

GLM-5.3 fits on 8x RTX Pro 6000s nvidia/GLM-5.3-NVFP4 - 54 decode tok/s - 251k context - 476 peak decode tok/s incoai/GLM-5.3-NVFP4 (DFlash2) - 93 decode tok/s - 123k context - 486 peak decode tok/s https://t.co/1wgf2kLaRI

Digest-only: incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × GLM-5.3 Flash

Sep 30, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
131072 NVFP4 vllm stock 15.0 tok/s —
— NVFP4 vllm stock 19.5 tok/s —
— NVFP4 vllm stock 31.7 tok/s —
— NVFP4 vllm stock 26.7 tok/s —
— NVFP4 vllm stock 31.7 tok/s —
122880 NVFP4 vllm stock 30.4 tok/s —
— NVFP4 vllm stock — 1350.0 tok/s
131072 NVFP4 vllm stock — 1580.0 tok/s

loading GLM-5.3 Flash on my 2x DGX Spark, 256GB of unified memory across two boxes, NVIDIA's NVFP4 weights split over the ConnectX cable on stock vLLM numbers so far: 320B total, 18B active, 204GB weights decode, MTP off: 15 tok/s, flat to 128K decode, MTP on (k=3): 19.5-31.7 tok/s code + math, MTP on: 26.7-31.7 tok/s 120K context, MTP on: 30.4 tok/s prefill, MTP off: 1,350-1,580 tok/s 128K prompt: first token in…

Digest-only: incomplete methodology (engine+version, measurement method).

DGX Spark × Qwen3.8-27B

Sep 30, 2026 · by @0xBakeer · third-party report

Context Quant Engine Decode Prefill
150K — — — 90.0 tok/s

Running Qwen3.8-27B (the dense model, not Qwen Flash) on a single DGX Spark. Thanks to optimized caching algorithms and the StairCut method, the ~150k context prefill takes just milliseconds, and decoding speeds hit around 50–90 tok/s depending on the task. In this clip, I recorded after 5 hours of continuous running, I had Qwen design a pagoda garden and keep iterating and implementing new ideas nonstop.

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

M5 Max 128GB × MiMo-V2.6-Flash

Sep 30, 2026 · by @Beamsters1

Context Quant Engine Decode Prefill
— 2.3bpw sushi v1.1 — 1132.0 tok/s
— 2.3bpw sushi v1.1 70.0 tok/s —

Sushi 🍣 v1.1 release with MiMo-V2.6-Flash! This 2.3bpw created a one-shot 3D racing game for browser, fully playable with sound effect. M5 Max 128GB - prefill 1132 tok/s, gen 70 tok/s - max 1m context. Get it here: https://github.com/beamivalice/sushi https://t.co/asXK1XDWGu

Full methodology stated.

Unspecified platform

Sep 30, 2026 · by @runinfrai · third-party report

Context Quant Engine Decode Prefill
1M FP8 — 670.0 tok/s —

we spent september rewriting the kernels behind GLM 5.3 Flash. major release is live on RunInfra today 670 tok/s on Vercel AI Gateway $0.11 per 1M input, $0.45 per 1M output, $0.03 per 1M cached. 1M token context. FP8, vendor-native release it now runs on AMD. same model, same API, more capacity behind it 99.7% cache hit rate over the last 24 hours. every hit and miss shows up in your dashboard, so you can track the…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, hardware state, measurement method).

RTX 5070 Ti × Qwen3.8-Flash-Next

Sep 30, 2026 · by @jurlycat

Context Quant Engine Decode Prefill
— IQ3_XXS strata 51.0 tok/s —
— IQ3_XXS strata — 1500.0 tok/s
— IQ3_XXS llama.cpp 23.0 tok/s —
— IQ3_XXS llama.cpp — 100.0 tok/s

A 125B MoE model just ran at 51 tok/s on a laptop: • RTX 5070 Ti, 12GB VRAM • Intel 275HX • 64GB DDR5 RAM • Gen4 SSD Strata splits Qwen3.8-Flash-Next across the GPU, RAM, and SSD. Same IQ3_XXS quant: • Strata: 51 tok/s generation, 1,500 tok/s prefill • llama.cpp: 23 tok/s generation, 100 tok/s prefill Models this large no longer need server-grade hardware.

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Sep 30, 2026 · by @Tech2Wild

Context Quant Engine Decode Prefill
— — — 35.0 tok/s —
— — — 69.0 tok/s —
— — — 71.0 tok/s —
— — — 65.0 tok/s —
— — — 101.0 tok/s —

⚡ GLM-5.3-Flash on just 2x DGX Spark: new default 🚀 We ported @majewskizby 4-Spark stack down to 2 nodes (TP2) 🔧 Same hardware, same benchmark, vs our old TP2 recipe: 📝 Prose 18 → 35 tok/s (+95%) 💻 Code 45 → 69 · JSON 43 → 71 · Math 46 → 65 🛠️ Tool calls 48 → 101 tok/s 📈 Multi-user throughput +19–62% (1–5 streams) ⏱️ Weights load in 44s (was ~11 min) 💾 560K-token KV pool: two full 262K contexts at once Two lanes…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x4 + RTX 5090 × GLM-5.3-Flash

Sep 30, 2026 · by @dangerm00se

Context Quant Engine Decode Prefill
79K — — — 5500.0 tok/s
79K — — 133.0 tok/s —

updated. i looked into the super fast tp4 recipes and applying the same lossy techniques could perhaps double prefill but decided the quality loss isn't worth it - this recipe is both fast and highly accurate. v1.1: GLM-5.3-Flash on 4 DGX Sparks + a 5090 now prefills 5.5K tok/s at 79K (+5% over v1.0) and decodes code at 133 tok/s on one stream (+5.5%), with byte-identical output. https://t.co/GEeYql40jo

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark x1 × Qwen3.8-Flash-Next

Sep 30, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— EXL3 3.05bpw exllamav2 custom fork with native engine path 102.6 tok/s —
— EXL3 3.05bpw exllamav2 custom fork with native engine path 71.7 tok/s —
— EXL3 3.05bpw exllamav2 custom fork with native engine path 66.2 tok/s —
8K EXL3 3.05bpw exllamav2 custom fork with native engine path 51.5 tok/s —
64K EXL3 3.05bpw exllamav2 custom fork with native engine path 67.4 tok/s —
240K EXL3 3.05bpw exllamav2 custom fork with native engine path 55.8 tok/s —

I have to retract my previous claim about the fastest Qwen3.8-Flash-Next setup on a single DGX Spark. Cruz shipped a custom ExLlamaV3 fork with a native engine path, and I ran the same model on the same Spark. The result: 102.6 tok/s on code 71.7 tok/s on prose 66.2 tok/s on JSON My previous setup and community recipes were around 48.6 tok/s. That’s more than 2x the speed on the right workload. The quant is EXL3 at…

Digest-only: incomplete methodology (hardware state).

RTX 3060 × Tiel-Coder-35B-A3B

Sep 30, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 — llama.cpp 43.0 tok/s —
262144 — llama.cpp — 570.0 tok/s

A 35B coding model running on a single RTX 3060 with 12GB VRAM is already wild. But the more interesting comparison is with Qwen3.8-27B. The setup: • RTX 3060 12GB • 16GB system RAM • Tiel-Coder-35B-A3B • 262K native context • ~43 tok/s decode • ~570 tok/s prefill The model is actually Ornith-1.5 re-quantized with a custom imatrix, with the Sharp template baked directly into the GGUF. Same base weights I’ve been…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 3060 × qwen3.8-27b

Sep 30, 2026 · by @seanhighness

Context Quant Engine Decode Prefill
96K — l0xre 40.0 tok/s —

8.65 gb version of qwen3.8-27b a Hybrid of 8.65GB. Full Qwen3.8-27B. Running on an RTX 3060. A new hybrid built from @Eschalabs W2 + IQ3, paired with a custom L0xRE runtime designed to let the model hit its full potential. • Full Escha W2 compatibility • New 8.65GB hybrid model • ~40 tok/s decode on a 3060 12GB • DFlash2 speculative decoding • 96K context • Real agentic workloads with Hermes 27B on a 12GB…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark × Qwen3.6-35B-A3B

Sep 30, 2026 · by @0x0SojalSec · third-party report

Context Quant Engine Decode Prefill
— — — 50.0 tok/s —

Open-cybersecurity model you can run locally. - Fine-tuned from Qwen3.6-35B-A3B - 35B MoE, 3B active parameters per token - 50 tok/s on a DGX Spark (publisher figure) - Built for SOC, DFIR, and detection engineering so alerts and artifacts do not have to leave the building. - https://t.co/9AqTsguuQU

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑