DiggTop ⌕

X local-model bench digest — Oct 10, 2026

As of Oct 10, 2026 — 29 posts collected from X. 74 measured runs catalogued, 29 kept to the digest only. Latest source post Oct 10, 2026 12:47 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.

RTX 4090 × Qwen3.8-Flash-Next

Oct 10, 2026 · by @MinLiBuilds

Context Quant Engine Decode Prefill
256K — strata 100.0 tok/s —
256K — strata 180.0 tok/s —

😋 部署上了,吃上好东西了,同事们直夸好用,相见恨晚。 能力比 降智的sol 强太多了,速度是 sol 的 8 倍。 Qwen3.8-Flash-Next,这次部署的是 UD-IQ4_XS 量化版本,借助 Strata 装进一张 48GB RTX 4090,跑满 256K 上下文和 6 路并发。 单人约 100 tok/s,6 人同时使用合计约 180tok/s,已经足够直接给团队当本地 Codex 后端。 他们都是非研发为主,使用之前,他们担心这东西难用,我跟他说比 Opus 4.6 还强,他们眼睛都直了。 不负众望,完美落地,接下来让他们实测一周,我负责把推理速度推到极限👍 下一个阶段,准备让我的 MUSE 接入我的本地模型,因为它们都装了 Tail Scale。 然后让我的 MUSE 注册一个 L 站。惊艳他们🔥

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

A100 40GB × Qwen3.8-Flash-Next

Oct 10, 2026 · by @posi_posi8

Context Quant Engine Decode Prefill
— — strata 116.9 tok/s —
— — strata — 2572.0 tok/s

Geminiに惰性で課金している人へ。 今話題の高速推論エンジン「Strata」を試してみませんか? Google AI ProのColab特典を使えば、A100 40GBで125BのQwen3.8-Flash-Next(約2bit量子化)を動かせます。 今回の実測は最大116.9 tok/s、Prefillは2,572 tok/s。 GPUを購入する前に、大規模LLMの推論速度を体感できます。 Colabを公開したので、GPUを選んで上から順に▶を押すだけ! Colabはこちら https://t.co/wBmbyryXk6 GPUの選び方は、右上の接続の横の三角→ランタイムのタイプを変更→A100→保存(スクショ参照)

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

RTX 3060 × Underdog-Saluki-27B

Oct 10, 2026 · by @DogukanUrker

Context Quant Engine Decode Prefill
170K IQ2-mix llama.cpp 22.0 tok/s —
170K IQ2-mix llama.cpp — 490.0 tok/s

running Underdog-Saluki-27B on a single RTX 3060. 12GB VRAM. (config below) it’s Qwen3.8-27b squeezed into 7.9GB (2-bit) and tuned to keep tool calling intact. dense, so every layer sits on the gpu. no expert offload. 170K context. ~22 tok/s decode, ~490 tok/s prefill. running my agentic coding benchmark on it now. will compare it against other Qwen3.8-27B quants next! config: llama-server -m…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

RTX 5070 12GB + RTX 5060 Ti 16GB × Qwen3.8-Flash-Next

Oct 10, 2026 · by @murasametech

Context Quant Engine Decode Prefill
64 IQ3_S strata v0.1.41 44.9 tok/s 55.1 tok/s

Qwen3.8-Flash-Nextを、RTX 5070 12GBとRTX 5060 Ti 16GBの2GPU構成で動かしました。 量子化はIQ3_S(GGUF)で、モデルファイルは合計約84GBあります。 VRAM合計28GBには収まらないサイズですが、MoE構造なのでexpert部分をメインメモリ(64GB)に置き、dense重みをGPUに載せる形で実行できます。 加えてMTPによる投機デコードで、先のトークンをまとめて予測して生成を速める構成です。 推論エンジンはStrata v0.1.41、入力64tokensの短文を3回測った結果は次のとおりです。 Decode:39.8〜52.2 tok/s(中央値44.9) Prefill:44.7〜65.6 tok/s GPUに載り切らないサイズのモデルでも、短い質問への応答なら十分実用的な速度が出ています。

Full methodology stated.

M3 Ultra × GLM-5.3-Flash

Oct 10, 2026 · by @Alisvolatprop12

Context Quant Engine Decode Prefill
— — — 40.6 tok/s —
32768 — — — 636.0 tok/s

M3U 맥스튜디오 유저들을 위한 큰 한 걸음 [ - GLM-5.3-Flash 디코드(M3 Ultra): 31.3 -> 40.6 tok/s (+30%). 융합 디코드 커널이 이제 M5뿐만 아니라 M3과 M4에서도 실행됩니다. - Qwen3.8-Flash-Next(전문가 오프로드 포함), M3 Ultra에서 32K 프리필: 396 -> 636 tok/s (+61%). ]

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

Unspecified platform × DeepSeek V4.1 Flash

Oct 10, 2026 · by @xrath · third-party report

Context Quant Engine Decode Prefill
— — — 450.0 tok/s —

1T 이하 모델 중 에이전트로 쓸만하다고 생각하는 건 DeepSeek V4.1 Flash 하나뿐인데 그마저도 max로 써야 믿을만하다. 지름길 놔두고 빙글빙글 돌아가는 느낌이 강하지만 결과물은 제대로 가져와서 만족. 이걸 450 tok/s 뽑아주는 업체가 있길래 $50 충전해서 써봤는데 코딩할 때 500 넘게 나옴. https://t.co/GzDDSqzK02

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

V100 32GB x2 × Qwen3.8-Flash-Next

Oct 10, 2026 · by @PhotogenicWeekE

Context Quant Engine Decode Prefill
— IQ3_S — 109.0 tok/s —
— IQ3_S — 99.0 tok/s —
— IQ3_S — 75.0 tok/s —
— IQ3_S — 81.0 tok/s —
23K IQ3_S — — 2000.0 tok/s

Strataさらに(笑) 完全バラックwww V100 32GB×2でQwen3.8-Flash-Next(IQ3_S) Code 109 / Logic 99 / Prose 75 / Japanese 81 tok/s Prefill: 約2,000 tok/s(23kトークン) 1枚の約1.3〜1.85倍。2枚目OCuLink接続の影響無し! https://t.co/GeZsFRg6jG

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

V100 × Strata

Oct 10, 2026 · by @PhotogenicWeekE

Context Quant Engine Decode Prefill
— — — 62.0 tok/s —
— — — 62.7 tok/s —
— — — 53.5 tok/s —
— — — 59.0 tok/s —
— — — 52.3 tok/s —
— — — 56.0 tok/s —
— — — 59.7 tok/s —
— — — 57.4 tok/s —
5900 — — — 985.0 tok/s
23K — — — 1356.0 tok/s

Strataベンチ結果(V100 1 枚、IQ3_S) ー none low Code 62.0 62.7 Logic 53.5 59.0 Prose 52.3 56.0 Japanese 59.7 57.4 prefill 5.9kで約 985 tok/s、23kで約 1,356 tok/s 意外と速い♪ https://t.co/fbujFSqenZ

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

RTX 3090 × Qwen3.8-27B

Oct 10, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
256K Q4 unsloth 79.6 tok/s —
— Q4 unsloth 84.4 tok/s —
196K Q4 unsloth 84.9 tok/s —
— Q4 vulkan 85.4 tok/s —

if you own a single rtx 3090, or any 24gb card, your local ai king is qwen 3.8 27b dense. nothing you can fit on that card comes close, and you can stop looking it fits at q4 in 16.8gb with room for a 256k window and vision, and with the mtp head that already ships inside the gguf the community numbers are wild: rtx 3090: 79.6 tok/s, overclocked, on unsloth's q4 rtx 3090 on a tool call workload: 84.4 tok/s with 97%…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

DGX Spark × DeepSeek 4.1 Flash

Oct 10, 2026 · by @voidwarriorchan · third-party report

Context Quant Engine Decode Prefill
— EXL3 — 70.0 tok/s —

GLM 5.3 Flash EXL3 AbliteratedをDGX Spark 2台でカリカリチューニングして40~70 tok/sくらい 長時間の稼働はDeepSeek 4.1 Flashよりもいい気がするので乗り換えた

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).

Unspecified platform × GLM-5.3-full

Oct 10, 2026 · by @mmastrac

Context Quant Engine Decode Prefill
— — kindlingai tp4 49.5 tok/s —
— — kindlingai tp4 30.4 tok/s —
8192 — kindlingai tp4 — 1292.0 tok/s
8192 — kindlingai tp4 — 55079.0 tok/s
— — kindlingai tp4 38.1 tok/s —
— — kindlingai tp4 50.9 tok/s —
— — kindlingai tp4 63.4 tok/s —

Excerpt from rigmark on glm 5.3 full on the kindlingai engine, tp4, beta testing now release soon. EXLR8 "home" quant. │ CODE 49.5 tok/s 48.4–49.9 ✓ 5/5 │ │ PROSE 30.4 tok/s 29.7–31.3 ✓ 5/5 │ │ 8K PREFILL cold 1,292 tok/s • cached replay 55,079 tok/s │ AGGREGATE C1 38.1 • C2 50.9 • C4 63.4 tok/s

Digest-only: hardware.name missing; incomplete methodology (hardware state).

RTX 3060 × Qwen3.8-35B-A3B

Oct 10, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
262144 Q4_K_M llama-server 50.0 tok/s —
262144 Q4_K_M llama-server — 550.0 tok/s

Someone is running Qwen3.8-35B-A3B on a single RTX 3060 with just 12GB of VRAM and 16GB of system RAM. And the reported numbers are surprisingly good. ~50 tok/s decode
~550 tok/s prefill
262K context window
12GB VRAM
16GB system RAM No multi-GPU setup. Just a consumer graphics card and a carefully configured local inference engine. The model is particularly interesting because it’s Qwen3.8 distilled into the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

Also in this window

The 12 most recent posts are shown in full above; 17 posts from the same window are listed here in short form, each with its source post.

How to read this page

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest sourced?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a row to reach the sourced tier?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑