X local-model bench digest — Oct 10, 2026
As of Oct 10, 2026 — 29 posts collected from X. 74 measured runs catalogued, 29 kept to the digest only. Latest source post Oct 10, 2026 12:47 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.
RTX 4090 × Qwen3.8-Flash-Next
Oct 10, 2026 · by @MinLiBuilds
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256K | — | strata | 100.0 tok/s | — |
| 256K | — | strata | 180.0 tok/s | — |
😋 部署上了,吃上好东西了,同事们直夸好用,相见恨晚。 能力比 降智的sol 强太多了,速度是 sol 的 8 倍。 Qwen3.8-Flash-Next,这次部署的是 UD-IQ4_XS 量化版本,借助 Strata 装进一张 48GB RTX 4090,跑满 256K 上下文和 6 路并发。 单人约 100 tok/s,6 人同时使用合计约 180tok/s,已经足够直接给团队当本地 Codex 后端。 他们都是非研发为主,使用之前,他们担心这东西难用,我跟他说比 Opus 4.6 还强,他们眼睛都直了。 不负众望,完美落地,接下来让他们实测一周,我负责把推理速度推到极限👍 下一个阶段,准备让我的 MUSE 接入我的本地模型,因为它们都装了 Tail Scale。 然后让我的 MUSE 注册一个 L 站。惊艳他们🔥
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
A100 40GB × Qwen3.8-Flash-Next
Oct 10, 2026 · by @posi_posi8
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | strata | 116.9 tok/s | — |
| — | — | strata | — | 2572.0 tok/s |
Geminiに惰性で課金している人へ。 今話題の高速推論エンジン「Strata」を試してみませんか? Google AI ProのColab特典を使えば、A100 40GBで125BのQwen3.8-Flash-Next(約2bit量子化)を動かせます。 今回の実測は最大116.9 tok/s、Prefillは2,572 tok/s。 GPUを購入する前に、大規模LLMの推論速度を体感できます。 Colabを公開したので、GPUを選んで上から順に▶を押すだけ! Colabはこちら https://t.co/wBmbyryXk6 GPUの選び方は、右上の接続の横の三角→ランタイムのタイプを変更→A100→保存(スクショ参照)
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 3060 × Underdog-Saluki-27B
Oct 10, 2026 · by @DogukanUrker
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 170K | IQ2-mix | llama.cpp | 22.0 tok/s | — |
| 170K | IQ2-mix | llama.cpp | — | 490.0 tok/s |
running Underdog-Saluki-27B on a single RTX 3060. 12GB VRAM. (config below) it’s Qwen3.8-27b squeezed into 7.9GB (2-bit) and tuned to keep tool calling intact. dense, so every layer sits on the gpu. no expert offload. 170K context. ~22 tok/s decode, ~490 tok/s prefill. running my agentic coding benchmark on it now. will compare it against other Qwen3.8-27B quants next! config: llama-server -m…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
RTX 5070 12GB + RTX 5060 Ti 16GB × Qwen3.8-Flash-Next
Oct 10, 2026 · by @murasametech
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 64 | IQ3_S | strata v0.1.41 | 44.9 tok/s | 55.1 tok/s |
Qwen3.8-Flash-Nextを、RTX 5070 12GBとRTX 5060 Ti 16GBの2GPU構成で動かしました。 量子化はIQ3_S(GGUF)で、モデルファイルは合計約84GBあります。 VRAM合計28GBには収まらないサイズですが、MoE構造なのでexpert部分をメインメモリ(64GB)に置き、dense重みをGPUに載せる形で実行できます。 加えてMTPによる投機デコードで、先のトークンをまとめて予測して生成を速める構成です。 推論エンジンはStrata v0.1.41、入力64tokensの短文を3回測った結果は次のとおりです。 Decode:39.8〜52.2 tok/s(中央値44.9) Prefill:44.7〜65.6 tok/s GPUに載り切らないサイズのモデルでも、短い質問への応答なら十分実用的な速度が出ています。
Full methodology stated.
M3 Ultra × GLM-5.3-Flash
Oct 10, 2026 · by @Alisvolatprop12
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 40.6 tok/s | — |
| 32768 | — | — | — | 636.0 tok/s |
M3U 맥스튜디오 유저들을 위한 큰 한 걸음 [ - GLM-5.3-Flash 디코드(M3 Ultra): 31.3 -> 40.6 tok/s (+30%). 융합 디코드 커널이 이제 M5뿐만 아니라 M3과 M4에서도 실행됩니다. - Qwen3.8-Flash-Next(전문가 오프로드 포함), M3 Ultra에서 32K 프리필: 396 -> 636 tok/s (+61%). ]
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).
Unspecified platform × DeepSeek V4.1 Flash
Oct 10, 2026 · by @xrath · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 450.0 tok/s | — |
1T 이하 모델 중 에이전트로 쓸만하다고 생각하는 건 DeepSeek V4.1 Flash 하나뿐인데 그마저도 max로 써야 믿을만하다. 지름길 놔두고 빙글빙글 돌아가는 느낌이 강하지만 결과물은 제대로 가져와서 만족. 이걸 450 tok/s 뽑아주는 업체가 있길래 $50 충전해서 써봤는데 코딩할 때 500 넘게 나옴. https://t.co/GzDDSqzK02
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
V100 32GB x2 × Qwen3.8-Flash-Next
Oct 10, 2026 · by @PhotogenicWeekE
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ3_S | — | 109.0 tok/s | — |
| — | IQ3_S | — | 99.0 tok/s | — |
| — | IQ3_S | — | 75.0 tok/s | — |
| — | IQ3_S | — | 81.0 tok/s | — |
| 23K | IQ3_S | — | — | 2000.0 tok/s |
Strataさらに(笑) 完全バラックwww V100 32GB×2でQwen3.8-Flash-Next(IQ3_S) Code 109 / Logic 99 / Prose 75 / Japanese 81 tok/s Prefill: 約2,000 tok/s(23kトークン) 1枚の約1.3〜1.85倍。2枚目OCuLink接続の影響無し! https://t.co/GeZsFRg6jG
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
V100 × Strata
Oct 10, 2026 · by @PhotogenicWeekE
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 62.0 tok/s | — |
| — | — | — | 62.7 tok/s | — |
| — | — | — | 53.5 tok/s | — |
| — | — | — | 59.0 tok/s | — |
| — | — | — | 52.3 tok/s | — |
| — | — | — | 56.0 tok/s | — |
| — | — | — | 59.7 tok/s | — |
| — | — | — | 57.4 tok/s | — |
| 5900 | — | — | — | 985.0 tok/s |
| 23K | — | — | — | 1356.0 tok/s |
Strataベンチ結果(V100 1 枚、IQ3_S) ー none low Code 62.0 62.7 Logic 53.5 59.0 Prose 52.3 56.0 Japanese 59.7 57.4 prefill 5.9kで約 985 tok/s、23kで約 1,356 tok/s 意外と速い♪ https://t.co/fbujFSqenZ
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
RTX 3090 × Qwen3.8-27B
Oct 10, 2026 · by @sudoingX
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256K | Q4 | unsloth | 79.6 tok/s | — |
| — | Q4 | unsloth | 84.4 tok/s | — |
| 196K | Q4 | unsloth | 84.9 tok/s | — |
| — | Q4 | vulkan | 85.4 tok/s | — |
if you own a single rtx 3090, or any 24gb card, your local ai king is qwen 3.8 27b dense. nothing you can fit on that card comes close, and you can stop looking it fits at q4 in 16.8gb with room for a 256k window and vision, and with the mtp head that already ships inside the gguf the community numbers are wild: rtx 3090: 79.6 tok/s, overclocked, on unsloth's q4 rtx 3090 on a tool call workload: 84.4 tok/s with 97%…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
DGX Spark × DeepSeek 4.1 Flash
Oct 10, 2026 · by @voidwarriorchan · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | — | 70.0 tok/s | — |
GLM 5.3 Flash EXL3 AbliteratedをDGX Spark 2台でカリカリチューニングして40~70 tok/sくらい 長時間の稼働はDeepSeek 4.1 Flashよりもいい気がするので乗り換えた
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).
Unspecified platform × GLM-5.3-full
Oct 10, 2026 · by @mmastrac
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | kindlingai tp4 | 49.5 tok/s | — |
| — | — | kindlingai tp4 | 30.4 tok/s | — |
| 8192 | — | kindlingai tp4 | — | 1292.0 tok/s |
| 8192 | — | kindlingai tp4 | — | 55079.0 tok/s |
| — | — | kindlingai tp4 | 38.1 tok/s | — |
| — | — | kindlingai tp4 | 50.9 tok/s | — |
| — | — | kindlingai tp4 | 63.4 tok/s | — |
Excerpt from rigmark on glm 5.3 full on the kindlingai engine, tp4, beta testing now release soon. EXLR8 "home" quant. │ CODE 49.5 tok/s 48.4–49.9 ✓ 5/5 │ │ PROSE 30.4 tok/s 29.7–31.3 ✓ 5/5 │ │ 8K PREFILL cold 1,292 tok/s • cached replay 55,079 tok/s │ AGGREGATE C1 38.1 • C2 50.9 • C4 63.4 tok/s
Digest-only: hardware.name missing; incomplete methodology (hardware state).
RTX 3060 × Qwen3.8-35B-A3B
Oct 10, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262144 | Q4_K_M | llama-server | 50.0 tok/s | — |
| 262144 | Q4_K_M | llama-server | — | 550.0 tok/s |
Someone is running Qwen3.8-35B-A3B on a single RTX 3060 with just 12GB of VRAM and 16GB of system RAM. And the reported numbers are surprisingly good. ~50 tok/s decode ~550 tok/s prefill 262K context window 12GB VRAM 16GB system RAM No multi-GPU setup. Just a consumer graphics card and a carefully configured local inference engine. The model is particularly interesting because it’s Qwen3.8 distilled into the…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
Also in this window
The 12 most recent posts are shown in full above; 17 posts from the same window are listed here in short form, each with its source post.
- RTX 3090 × Qwen3.8-27B — 133.0 tok/s decode (3 measurements) · @Oluwaphilemon1
- A100 — 116.9 tok/s decode (2 measurements) · @posi_posi8
- RTX PRO 6000 Blackwell × GLM-5.3-Flash — 110.0 tok/s decode (3 measurements) · @Nateemerson
- RTX PRO 6000 Blackwell × GLM-5.3-Flash — 110.0 tok/s decode (3 measurements) · @Nateemerson
- RTX 3060 × Bonsai 2 27B — 50.1 tok/s decode (1 measurement) · @sudoingX
- Unspecified platform × Qwen3.8-Flash-Next — 70.0 tok/s decode (1 measurement) · @mark_l_watson
- Intel Xeon CPU x4 × TwIL-LM3-Pro — 10.5 tok/s decode (3 measurements) · @thewebAI
- Unspecified platform × deepseek v4.1 flash — 217.0 tok/s decode (2 measurements) · @yuhasbeentaken
- MI355X x8 × Kimi K3 — 216.0 tok/s decode (1 measurement) · @lightseekorg
- Strix Halo × Qwen3.8 — 556.0 tok/s decode (1 measurement) · @no_stp_on_snek
- M1 Max × Qwen3-4B — 62.0 tok/s decode (1 measurement) · @LittleBit_llm
- RTX 3070 × GLM-5.3-Flash — 3.2 tok/s decode (1 measurement) · @bountyAIhunter
- M5 Max × Qwen3.6-35B-A3B — 205.0 tok/s decode (3 measurements) · @jundotkim
- DGX Spark x2 × DeepSeek-V4.1-Flash — 142.0 tok/s decode (3 measurements) · @bertholomusai
- Unspecified platform × Qwen3.8-27B — 264.0 tok/s decode (1 measurement) · @0xBakeer
- RTX 5090 × Qwen3.8-27B — 366.0 tok/s decode (5 measurements) · @niklaslenz_ai
- Strix Halo × qwen3.8-flash-next — 50.0 tok/s decode (1 measurement) · @drdanielbender
How to read this page
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @MinLiBuilds — Oct 10, 2026
- @posi_posi8 — Oct 10, 2026
- @DogukanUrker — Oct 10, 2026
- @murasametech — Oct 10, 2026
- @Alisvolatprop12 — Oct 10, 2026
- @xrath — Oct 10, 2026
- @PhotogenicWeekE — Oct 10, 2026
- @PhotogenicWeekE — Oct 10, 2026
- @sudoingX — Oct 10, 2026
- @voidwarriorchan — Oct 10, 2026
- @mmastrac — Oct 10, 2026
- @Oluwaphilemon1 — Oct 10, 2026
- @Oluwaphilemon1 — Oct 10, 2026
- @posi_posi8 — Oct 10, 2026
- @Nateemerson — Oct 10, 2026
- @Nateemerson — Oct 10, 2026
- @sudoingX — Oct 10, 2026
- @mark_l_watson — Oct 9, 2026
- @thewebAI — Oct 9, 2026
- @yuhasbeentaken — Oct 9, 2026
- @lightseekorg — Oct 9, 2026
- @no_stp_on_snek — Oct 9, 2026
- @LittleBit_llm — Oct 9, 2026
- @bountyAIhunter — Oct 9, 2026
- @jundotkim — Oct 9, 2026
- @bertholomusai — Oct 9, 2026
- @0xBakeer — Oct 9, 2026
- @niklaslenz_ai — Oct 9, 2026
- @drdanielbender — Oct 9, 2026
FAQ
Are the numbers in this digest sourced?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a row to reach the sourced tier?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.