X local-model bench digest — Oct 8, 2026
As of Oct 8, 2026 — 25 posts collected from X. 69 measured runs catalogued, 1 published to the board, 24 kept to the digest only. Latest source post Oct 8, 2026 10:28 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.
RTX 4090 × Qwen3.8-Flash-Next
Oct 8, 2026 · by @MinLiBuilds
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | Q2_0 | strata v0.1.40.3 | 288.0 tok/s | — |
125B的巨物Qwen3.8-Flash-Next,哪个版本塞入12GB的显存里效果最好? 为什么速度最快288tok/s的Q2,答题成绩和速度反而不是最佳的? Strata作者@coldniko与我讨论,决定详细测试4个Q2版本的实际表现,以选出最佳推荐: Swift 1.5 Q2_0 Swift 1.5 IQ2_XS ISTA Q2_0 ISTA IQ2_XS 48GB RTX 4090,统一使用 Strata v0.1.40.3。 质量任务中的平均 decode,最高接近 288 tok/s。 确实飞起来了。 结论: 1. 我推荐ISTA IQ2_XS:答对最多,PPL最低,平均每题耗时也最短。 2. 低位量化,除了准确率和 tok/s,答题时会不会陷入循环思考也是重要指标,需要尽量避免。 详细结果👇 线程 1/5
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).
RTX 5070 × Qwen3.8-Flash-Next
Oct 8, 2026 · by @jinglian
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | strata | 94.0 tok/s | — |
| 128K | — | strata | 76.0 tok/s | — |
🔥 一个开源新框架,直接让顶配 Mac 沦为时代的眼泪。 闲鱼上的 M5 Ultra Mac Studio(36核/256GB)价格已经开始崩盘。 老黄的 DGX Spark 还在呲呲呲涨价 现在本地跑千亿大模型,一张中端 N 卡就彻底够用了。 ━━━━━ 💡 核心架构:Strata 的暴力调度 它的逻辑非常残暴,直接把 GPU、CPU、RAM、SSD 拧成一股绳: 🔹 常用专家层常驻显存 🔹 其余专家塞进系统内存 🔹 显存没命中?瞬间切给 CPU 兜底 实践哥的实测数据极其夸张: 一张 12G 显存的 RTX 5070 搭配 64G 内存,跑 125B 的
Qwen3.8-Flash-Next: 👉 基础推理速度飙到 94 tok/s 👉 哪怕拉满 128K 超长上下文,依然稳在 76 tok/s ━━━━━ 🤖 更离谱的是:整个部署压根就不是人干的 全流程由 560B 的Ling-3.1-flash接管…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).
GPU with 16GB VRAM × Qwen3.8-Flash-Next
Oct 8, 2026 · by @FiniYang
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ2_XS | — | 32.0 tok/s | — |
Strata 绝对是个好项目啊 本地居然能部署 Qwen3.8-Flash-Next 我 16G显存 + 32G内存可以玩 Coder IQ1_M 和 完整版IQ2_XS IQ2_XS 准确性明显提升 但是速度很慢(约 32 tok/s) 内存压力极大 我感觉我要去升级下内存了 🤔 项目地址 ⬇️ https://t.co/cyBBo5wpTI
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark × deepseek v4 flash
Oct 8, 2026 · by @sudoingX · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | fp4 | — | 16.5 tok/s | — |
| — | fp4 | — | 42.0 tok/s | — |
nvidia is putting dgx spark class silicon in windows laptops and launching on oct 16, i already know how this chip runs 1 petaflop fp4 blackwell gpu, 6,144 cuda cores, same count as my dgx spark 20 core grace cpu, up to 128gb unified memory a cheaper 5,120 core version with 24 or 32gb the number that decides your tok/s is bandwidth, nvidia hasn't published it yet, my dgx spark measures 273 gb/s on gb10, that's…
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 5070 × Qwen3.8-Flash-Next
Oct 8, 2026 · by @Oluwaphilemon1 · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | Q2_0 | — | 94.0 tok/s | — |
| — | Q2_0 | — | 53.0 tok/s | — |
Qwen3.8-Flash-Next went from “massive model that needs a server” to something people are running on gaming PCs in roughly two weeks. I wanted to understand what actually changed. The model itself didn’t suddenly become small. The inference stack changed. Qwen3.8-Flash-Next is a huge MoE configuration with roughly 180B total parameters, but only around 6B parameters are active per token. That changes everything…
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 4060 × Qwen3.8-27B
Oct 8, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 65536 | IQ4_XS | unsloth | — | 150.0 tok/s |
| 65536 | IQ4_XS | unsloth | 5.0 tok/s | — |
Qwen3.8-27B running on an RTX 4060 with 8GB of VRAM is a good example of why local AI hardware limits are becoming less straightforward. A 27B model on an 8GB card sounds impossible until you look at how the inference stack is actually using memory. The setup here is: RTX 4060, 8GB VRAM Qwen3.8-27B Unsloth IQ4_XS quant 14.6GB model on disk 64K context Native MTP ~150 tok/s prefill ~5 tok/s decode 25 layers offloaded…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
RTX 3090 × Qwen3.8-27B
Oct 8, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262144 | — | llama.cpp custom fork | 69.9 tok/s | — |
| 262144 | — | llama.cpp custom fork | — | 1124.0 tok/s |
| 82K | — | llama.cpp custom fork | 38.0 tok/s | — |
| 60K | — | llama.cpp custom fork | 27.9 tok/s | — |
Qwen3.8-27B running at nearly 70 tok/s on a 12GB GPU with a 262K context window sounds like one of those numbers that needs an asterisk. This time, the asterisk is the engineering. The setup is Mirai S Qwen3.8-27B, a 2.4-bit compressed version that weighs roughly 11GB. The model has been converted into GGUF form for a custom llama.cpp fork, with the compressed weights preserved rather than being re-quantized. The…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
DGX Spark × CYBER-FROST-3.8
Oct 8, 2026 · by @yume_arasaki
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 400 | EXL3 | exllamav2 047ce72 | 60.7 tok/s | — |
| 400 | EXL3 | exllamav2 047ce72 | 42.7 tok/s | — |
| — | EXL3 | exllamav2 047ce72 | 60.3 tok/s | — |
| — | EXL3 | exllamav2 047ce72 | 39.9 tok/s | — |
| — | EXL3 | exllamav2 047ce72 | 47.3 tok/s | — |
Yesterday I tested @Blackfrost_AI CYBER-FROST-3.8 , an uncensored 180B version of Qwen 3.8 Flash Next, on a 6 hour test on my own DGX Spark on its cybersecurity capabilities. Today is part 2, I show you the cleanest way to run it on one DGX Spark. The fastest way I have found to run this pack is not a new engine. It is @ViC305 DGX Spark recipe: his exllamav3 fork with mixed-K cooperative decode kernels, compiled for…
Full methodology stated (engine, quantization, context, hardware state, method) — 1 added to the benchmark board (1 of 5 measured rows queued; the rest are published here only).
AMD Strix Halo laptop × Qwen3.8-Flash-Next
Oct 7, 2026 · by @EugeneSmarts
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | Q4 | unsloth | 27.0 tok/s | — |
| 262K | Q4 | unsloth | 55.0 tok/s | — |
| 262K | Q4 | unsloth | — | 1400.0 tok/s |
A 262k-token prompt used to require an enterprise cloud API cluster, but an AMD Strix Halo laptop just sustained 27 tokens per second running Qwen3.8-Flash-Next locally. Strata just rolled out experimental Strix Halo support, scaling all the way to a 1M context window without a massive decode drop on Unsloth Q4 and GSQ-RCO weights. The telemetry coming out of testing: • Bosgame M5 in Balanced mode holds steady…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
Gb300 × deepseek v4.1 flash
Oct 7, 2026 · by @sudoingX · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | nvfp4 | vllm v4.1 | 155.0 tok/s | — |
| — | nvfp4 | vllm v4.1 | 599.0 tok/s | — |
| — | nvfp4 | vllm v4.1 | 34.0 tok/s | — |
vllm got 5.3x more throughput out of deepseek v4.1 flash in 3 weeks on same gb300s, only the software changed 5.3x throughput at 155 tok/s per user, 27k to 144k tok/s per chip 1.9x at the top end, 311 to 599 tok/s per user nvfp4 kv cache, 45% smaller than fp8 day 0 build on sep 11 vs oct 2, semianalysis numbers my two dgx sparks run v4.1 flash at about 34 tok/s on code today with speculative decoding on, so this…
Digest-only: confidence 0.35 below 0.60; incomplete methodology (context length, hardware state, measurement method).
RTX 3090 x2 × Claude Code Opus 5.5
Oct 7, 2026 · by @IbrahimSait_
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | codex | 60.5 tok/s | — |
Wait… 60.5 tok/s on this rig?! After getting GLM-5.3-Flash running with Codex, I switched to Claude Code with Opus 5.5 to optimize it. Same setup: 2× RTX 3090 · 32GB DDR4 · single NVMe From ~8.5–9.4 to 53.8 tok/s, peaking at 60.5 Using a lossy fast mode that skips some experts. Tool calling still passes!
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x2 × Qwen3.8 Flash Next
Oct 7, 2026 · by @MiaAI_lab
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1M | NVFP4 | tensorfold | 67.4 tok/s | 2500.0 tok/s |
| 1M | NVFP4 | tensorfold | 152.0 tok/s | 2500.0 tok/s |
| 1M | NVFP4 | tensorfold | 390.0 tok/s | 2500.0 tok/s |
| 1M | Int4 AutoRound | tensorfold | 89.0 tok/s | 3300.0 tok/s |
| 1M | Int4 AutoRound | tensorfold | 220.0 tok/s | 3300.0 tok/s |
| 1M | Int4 AutoRound | tensorfold | 483.0 tok/s | 2700.0 tok/s |
Run Qwen3.8 Flash Next with TensorFold on your 2x @NVIDIAAI DGX Sparks ⚡️ One model, two quants. Nvidia NVFP4: - 1M context, 3.9 KV cache - 67.4 tok/s on prose, single stream - 152 tok/s on prose, 4 concurrent streams - 390 tok/s on prose, 16 concurrent streams - Prefill 2500-2600 across the board Int4 AutoRound: - 1M context. 4.6M KV cache - 89 tok/s on prose, single stream - 220 tok/s on prose, 4 concurrent…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
Also in this window
The 12 most recent posts are shown in full above; 13 posts from the same window are listed here in short form, each with its source post.
- M5 Ultra x2 × GLM-5.3-Flash — 114.0 tok/s decode (2 measurements) · @0x0SojalSec
- Unspecified platform × mistral large — 68.0 tok/s decode (1 measurement) · @sudoingX
- M5 Ultra x2 × GLM 5.3 Flash — 114.0 tok/s decode (1 measurement) · @MiaAI_lab
- 2-GPU × Strata-iq3_s — 45.0 tok/s decode (7 measurements) · @tamanekokoro
- DGX Spark x4 × GLM-5.3 743B — 63.0 tok/s decode (3 measurements) · @Tech2Wild
- RTX 3090 × Mirai S Qwen3.8-27B — 69.9 tok/s decode (5 measurements) · @superalesha
- RTX 6000 × GLM-5.3 — 188.0 tok/s decode (1 measurement) · @kachowtowmater
- RTX 3090 × Qwen3.8-27B — 382.0 tok/s decode (2 measurements) · @Oluwaphilemon1
- DGX Spark x2 × DeepSeek-V4.1-Flash-EXL3 — 97.4 tok/s decode (4 measurements) · @sfxnz
- RTX 3090 x2 × GLM-5.3-Flash — 8.9 tok/s decode (1 measurement) · @IbrahimSait_
- RTX 3060 Ti 8GB × Bonsai 2 27B — 67.3 tok/s decode (5 measurements) · @sudoingX
- RTX 3090 — 21.0 tok/s decode (2 measurements) · @0xSero
- Arc Pro B70 × Qwen3.8 Flash Next — 36.0 tok/s decode (3 measurements) · @VRAMCalculator
How to read this page
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @MinLiBuilds — Oct 8, 2026
- @jinglian — Oct 8, 2026
- @FiniYang — Oct 8, 2026
- @sudoingX — Oct 8, 2026
- @Oluwaphilemon1 — Oct 8, 2026
- @Oluwaphilemon1 — Oct 8, 2026
- @Oluwaphilemon1 — Oct 8, 2026
- @yume_arasaki — Oct 8, 2026
- @EugeneSmarts — Oct 7, 2026
- @sudoingX — Oct 7, 2026
- @IbrahimSait_ — Oct 7, 2026
- @MiaAI_lab — Oct 7, 2026
- @0x0SojalSec — Oct 7, 2026
- @sudoingX — Oct 7, 2026
- @MiaAI_lab — Oct 7, 2026
- @tamanekokoro — Oct 7, 2026
- @Tech2Wild — Oct 7, 2026
- @superalesha — Oct 7, 2026
- @kachowtowmater — Oct 7, 2026
- @Oluwaphilemon1 — Oct 7, 2026
- @sfxnz — Oct 7, 2026
- @IbrahimSait_ — Oct 7, 2026
- @sudoingX — Oct 7, 2026
- @0xSero — Oct 7, 2026
- @VRAMCalculator — Oct 7, 2026
FAQ
Are the numbers in this digest sourced?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a row to reach the sourced tier?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.