DiggTop

X local-model bench digest — Sep 17, 2026

As of Sep 17, 2026 — 12 posts collected from X. 31 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Sep 17, 2026 12:58 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 18 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

RTX 3090 × Qwen3.8-Flash-Next

Sep 17, 2026 · by @mr_r0b0t

Context Quant Engine Decode Prefill
262144 EXL3 2.50bpw exllamav2 v1.5.0-compatible build with MoE CPU-offload sup… 38.6 tok/s
262144 EXL3 2.50bpw exllamav2 v1.5.0-compatible build with MoE CPU-offload sup… 27.8 tok/s
175K EXL3 2.50bpw exllamav2 v1.5.0-compatible build with MoE CPU-offload sup… 664.0 tok/s
175K EXL3 2.50bpw exllamav2 v1.5.0-compatible build with MoE CPU-offload sup… 20.9 tok/s
EXL3 2.50bpw exllamav2 v1.5.0-compatible build with MoE CPU-offload sup… 51.2 tok/s

Want to run Qwen3.8-Flash-Next on your RTX 3090? r0b0tlab/Qwen3.8-Flash-Next-EXL3-2.50bpw -Self-quantized, MTP + vision retained 👀 -38.6 tok/s in the speed sweep, 262K ctx -Perfect 20/20 score on the multi-turn BFCL subset Weights, runtime & evals: https://github.com/r0b0tlab/qwen38-flashnext-exl3 https://t.co/5JBroj0uTD

Full methodology stated (engine, quantization, context, hardware state, method) — 2 added to the benchmark board (2 of 5 measured rows queued; the rest are published here only).

DGX Spark x2 × GLM-5.3-Flash-NVFP4

Sep 17, 2026 · by @tenhkspark

Context Quant Engine Decode Prefill
NVFP4 vllm 35.1 tok/s

DGX Spark 2台でGLM-5.3-Flash-NVFP4が35.09 tok/s。1ノード対・単発・散文64問・温度0・thinking無効。自分の変更はBF16密行列のW4A16再量子化、NCCLのRDMA化、MTPドラフトK=2。品質はパープレキシティ5.1%悪化。重みと計測表: https://huggingface.co/tenhkspark/GLM-5.3-Flash-NVFP4-Wabi

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length).

RTX 5090 x1 + DGX Spark x4 × DeepSeek-V4.1-Flash

Sep 17, 2026 · by @LotusDecoder

Context Quant Engine Decode Prefill
4100.0 tok/s
5700.0 tok/s
74.0 tok/s
85.0 tok/s
168.0 tok/s

早上还在构思, 一张小显存高算力显卡,拖大显存低算力一体机, 原来已经有人做出来。 效益还很好。 --- 核心创新是把 DeepSeek-V4.1-Flash 这种超大 MoE 模型按照计算特征“拆开跑”: 用一张 RTX 5090 负责 Attention、KV Cache、投机解码器以及其他顺序执行部分, 再用 4 台 DGX Spark 专门存放和计算 MoE 的 routed experts。 这样既利用了 5090 强大的计算吞吐,又利用了 DGX Spark 的 128GB 统一内存容量。 最终这套 5090 + 4×Spark 的混合系统能够真正提供 1,048,576 tokens 上下文, 冷 Prefill 达约 4,100–5,700 tok/s, 单路 Decode 约 74–85 tok/s, 6 路并发约 168 tok/s 总吞吐。 更重要的是,利用 Prefix Cache,一个已经处理过的…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 3090 Founders Edition × Qwen3.8-27B

Sep 17, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
8192 EXL3 4.0bpw exllamav2 152.4 tok/s

Qwen3.8-27B is now pushing 152 tok/s on a single RTX 3090. The setup is: Qwen3.8-27B 4.0 bpw EXL3 quant DFlash2-EXL3 Optimized ExLlamaV3 runtime RTX 3090 Founders Edition Measured result: 152.4 tok/s at 8K context. And the part that makes this even more interesting is the memory footprint. Peak VRAM was 22.7 GiB. The model also supports up to 262K context capacity. So we’re talking about a 27B model running locally…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

Unspecified platform × Qwen3.8-27B

Sep 17, 2026 · by @TeksEdge · third-party report

Context Quant Engine Decode Prefill
Q4_K_XL 845.0 tok/s
Q4_K_XL 73.8 tok/s
Q4_K_XL 95.1 tok/s

Long enough horizon. Qwen3.8-27B's best community benchmark score may be (30 days). Local AI in production. A Local AI user on Reddit says they ran Qwen3.8-27B UD-Q4_K_XL every day for a month for coding, long agent sessions, and overnight runs lasting 9+ hours. Their reported averages: ⚡ 845 tok/s prefill 🚀 73.8 tok/s generation 🔥 95.1 tok/s peak 🧠 ~48% MTP acceptance But the performance numbers aren’t the…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

DGX Spark × qwen

Sep 17, 2026 · by @sudoingX · third-party report

Context Quant Engine Decode Prefill
256K fp8 45.0 tok/s

here is the video of qwen 3.8 flash next autonomously building a visually striking website, start to finish, 29 minutes at 9.7x speed. the stack: official fp8 on 2x dgx spark, tensor parallel over one cable, 256k context loaded, 45 tok/s sustained with mtp on, hermes agent driving it from my laptop. the input was one spec file. it came back with this and started a server on my tailnet so i could open it from the…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, hardware state, measurement method).

RX 7900 XT × Qwen3.8-27B

Sep 17, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
65536 llama.cpp 30.8 tok/s
65536 llama.cpp 54.0 tok/s
65536 llama.cpp 54.7 tok/s
65536 llama.cpp 53.1 tok/s
65536 llama.cpp 65.0 tok/s
65536 llama.cpp 34.9 tok/s

amd users, rejoice. someone put an rx 7900 xt through the same a/b as everyone else, on omarchy, rocm 7.2, headless, and the gaming card flies: 30.8 tok/s stock to 54.0 tok/s with one llama.cpp flag on, +75%, qwen 3.8 27b dense at 64k context on 20gb of vram with 2.5gb to spare. he swept it properly. draft depth 2, 3 and 4 land at 54.0, 54.7 and 53.1 tok/s, code climbs to 65.0 tok/s at depth 4 while prose falls to…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).

RTX 3060 × Qwen3.8-35B-A3B

Sep 17, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 Q4_K_M llama.cpp 50.0 tok/s
262144 Q4_K_M llama.cpp 550.0 tok/s

A 12GB RTX 3060 running a 35B model at around 50 tok/s is already pretty ridiculous. But the more interesting part isn’t the speed. It’s what model is sitting behind it. Empero’s Qwen3.8-35B-A3B is running on: RTX 3060 12GB 16GB system RAM 262K context ~550 tok/s prefill ~50 tok/s decode And this isn’t just a tiny 35B model squeezed into memory with a huge performance compromise. It’s Qwen3.8 distilled into the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash-EXL3-2.9bpw

Sep 16, 2026 · by @no_stp_on_snek

Context Quant Engine Decode Prefill
61119 EXL3 2.9bpw deepseek harness 36.8 tok/s
50688 deepseek harness 28.2 tok/s

Given that WOW is in the news, Deepseek 4.1 flash (left) vs GLM 5.3 Flash (right) making Jimothy's penthouse in "minecraft" (both quantized to fit on 2 sparks). Deepseek looks better. I couldn't even zoom out enough on GLM to get a better view. Setup (same for both) Same prompt, DeepSeek Harness, high think 2× DGX Spark, TP=2, EXL3 0 retries on the completion turn What was served DeepSeek-V4.1-Flash…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).

RTX 4080 × Qwen3.8-27B

Sep 16, 2026 · by @yume_arasaki · third-party report

Context Quant Engine Decode Prefill
Q3_K_M 46.0 tok/s

Hugging Face is not banning uncensored models. I checked, but signs of raising the drawbridge for plebs has begun. The goal: Only insiders have access to unrestricted AI, thats why I decided to put this list together The local scene has been saying this out loud. Your weights live on one company's disk. One policy shift and the page is gone. I wrote about abliteration last week. It is not a jailbreak prompt. It is…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 4060 × Qwen3.8-27B

Sep 16, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
65536 IQ4_XS unsloth 150.0 tok/s
65536 IQ4_XS unsloth 5.0 tok/s

Qwen3.8-27B running on an RTX 4060 with 8GB VRAM is one of those local AI setups that makes the spec sheet look misleading. The model itself is a 27B dense model. The GPU has only 8GB. Yet the setup reaches a 64K context window. The trick is the quantization. Unsloth’s new IQ4_XS build brings the model down to roughly 14.6GB on disk. That still doesn’t fit inside 8GB, obviously. So the inference workload is split…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 3060 × Qwen3.8-35B-A3B

Sep 16, 2026 · by @DogukanUrker

Context Quant Engine Decode Prefill
262144 Q4_K_M llama.cpp 50.0 tok/s
262144 Q4_K_M llama.cpp 550.0 tok/s

running Empero's Qwen3.8-35B-A3B on a single RTX 3060. 12GB VRAM, 16GB RAM. (config below) full 262K context. ~50 tok/s decode, ~550 tok/s prefill. it's Qwen3.8 distilled into the Qwen3.6-35B-A3B architecture. Ornith-1.5 is the engine behind my Hermes and Pi agents right now, so the real question is whether this can replace it. next: running it through my agentic coding benchmark against Ornith-1.5-35B, Qwen3.8-27B…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.