X local-model bench digest — Oct 11, 2026
As of Oct 11, 2026 — 27 posts collected from X. 66 measured runs catalogued, 1 published to the board, 26 kept to the digest only. Latest source post Oct 11, 2026 09:01 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Only posts that state the full methodology (engine build, quantization, context length, hardware state, measurement method) reach the benchmark board; the rest stay here as leads.
DGX Spark × Qwen3.8-27B
Oct 11, 2026 · by @ashxhart · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | — | 125.0 tok/s | — |
TensorFold 1.0.7 has landed with a major overhaul to Living Weights. • Each fact you teach through /learn becomes about ten new neurons inside the model's own weights, learned about 4x faster, with a live 3D view of what it has learned. Macs and DGX Spark. • Structured output and required tool calls on the native server, with drafting still on. • Qwen3.8-27B serves concurrent requests, and sampled replies draft too…
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 5090 x1 + RTX 5060 Ti × Qwen3.8 Flash-Next
Oct 11, 2026 · by @ai_hakase_
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ3_XXS | basalt | 665.0 tok/s | — |
| — | IQ3_XXS | basalt | 354.0 tok/s | — |
【Blackwell専用の極限カスタム推論エンジン「Basalt」登場!】 既存のStrataを大幅に再構築し、Blackwell世代のGPUに特化したLLM推論エンジン「Basalt」が公開されました!デコードおよびプレフィルカーネルの大部分を書き換え、旧世代Nvidia GPUやAMD、Intel環境のサポートをあえて切り捨てることで極限的な高速化を実現しています。 RTX 5090とRTX 5060 TiのデュアルGPU環境において、Qwen3.8 Flash-Nextモデル(IQ3_XXS量子化)を使い、構造化出力で665 tok/s(Strataの約2.6倍)、通常文章生成で354 tok/sという高いスループットを実測しているんです!また、単一の
.basaltファイルへのウェイト等再パッキングや、ファインチューニングされたMTP(Multi-Token…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 5070 x1 + RTX 5060 Ti × Qwen3.8-Flash-Next
Oct 11, 2026 · by @murasametech
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ3_S | strata | 41.5 tok/s | — |
ローカルLLMをAIエージェントで使ってみて、ベンチマークの数字と実用性は別物だと痛感しました。 構成はRTX 5070 12GB+RTX 5060 Ti 16GB、Qwen3.8-Flash-Next(IQ3_S)、推論エンジンStrata、エージェントはOpenCodeです。 課題は総務省統計局「家計調査」のCSVを集計するPythonツールの作成。CSVはShift_JISで、見出しが複数行にまたがる表です。 AIはiconvで文字コードを変換し、awkで列や見出しを確かめるところから始めました。ところが30分たってもその調査段階から先に進まず、ファイルを1つも作らないまま中断しました。 生成速度は約34〜49 tok/sで、数字だけ見れば十分な速さです。 それでもコマンドを1回実行するたびに、次の応答まで1〜2分待たされる状態でした。…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
M4 Mac mini + RTX 3090 × Qwen3.8-Flash-Next
Oct 11, 2026 · by @wei_wang
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 200K | 2 bit | tinygrad | 80.0 tok/s | — |
我手边也有一台 M4 Mac mini,看到这个组合有点心动。 16GB 的 M4 Mac mini 用 USB-C 外接一张 RTX 3090,跑 Qwen3.8-Flash-Next: 1. 大部分活跃权重放在 3090 上,其余交给 Strata,用 12GB 的 Mac 内存和 256GB 的内置存储 2. 2 bit 量化,200K KV cache 3. decode 约 80 tok/s,整套 $2,800 外接显卡走的是 tinygrad 的方案。2 bit 的质量要打个问号,但 $2,800 能跑 Flash-Next,门槛又低了一截。
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
DGX Spark x2 × RED-SNOW-5.3-FLASH
Oct 11, 2026 · by @ViC305
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 | exllamav2 | 53.0 tok/s | — |
53 tok/s! 🚀Phenomenal work by @justinn_builds getting the GLM-5.3-Flash cyber fine-tuned ablit RED-SNOW-5.3-FLASH EXL3 collab with @Blackfrost_AI to 53tps across two DGX Sparks!! Bravo my brother, just another example of the community doing its thing and working its magic with my EXL3 quants!! Grateful for all the time invested 🙏 Recipe: https://t.co/AeQOFX13oJ Quant: https://t.co/xA7rZFe3Z3
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 5070 12GB + RTX 5060 Ti 16GB × Qwen3.8-Flash-Next
Oct 11, 2026 · by @murasametech
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 14K | IQ3_S | strata | — | 151.5 tok/s |
| 14K | IQ3_S | strata | 34.3 tok/s | — |
ローカルLLMでは、入力が長くなると待ち時間の中身が大きく変わります。 Qwen3.8-Flash-Next(IQ3_S)をRTX 5070 12GB+RTX 5060 Ti 16GBの2GPU構成、推論エンジンStrataで動かし、約14,000tokensの長文を入力してみました。 Prefill:94.78秒(151.5 tok/s) Decode:34.3 tok/s 回答完了まで:98.4秒 約98秒のうち約95秒が、入力を読み込むPrefillの時間でした。 一方、入力64tokensの短文では、Prefillは1秒前後で、待ち時間の大半は回答を生成するDecode(5〜6秒)です。 短い質問ではDecode速度が体感を左右しますが、長いコードや文書を渡す用途ではPrefillが待ち時間のほとんどを占めます。…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
A100 × Strata
Oct 11, 2026 · by @posi_posi8
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | IQ3_S | — | — | 2603.0 tok/s |
| — | IQ3_S | — | 103.0 tok/s | — |
Google ProについてくるA100を使ってStrataを動かすスクリプトについて、量子化度合いに関するリプライが多かったので、モデルを選択できるように更新した。 IQ3_Sでも、長文入力時はPrefill 2,603 tok/s、Decode 103 tok/s(64トークン生成)で、IQ2_XSと大差なかった。 https://colab.research.google.com/drive/1eop4zr2ofgRJDfdv6nZL4Lyo7vU_uPeU?usp=sharing https://t.co/zh2H7XwGAM
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
DGX Spark × Nemotron 3 Super
Oct 11, 2026 · by @ScottLeimroth
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | NVFP4 | vllm | 15.0 tok/s | — |
| — | NVFP4 | vllm | 26.0 tok/s | — |
I ran NVIDIA's Nemotron 3 Super (120B, 12B active) in NVFP4 on a single DGX Spark using NVIDIA's own vLLM recipe: Agent test: 56/60. The hosted full-precision version scored 57-58. Speed: 15 tok/s plain, 26 tok/s with its built-in MTP drafter. But with MTP on, the agent score dropped to 53. Long context held: three ~120k-token requests, every needle found. JSON output: 19/20 valid with thinking off, only 11/20 with…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).
RTX 5070
Oct 10, 2026 · by @jun1228909
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | Q2_0 | strata v0.1.42 | 85.6 tok/s | — |
小規模とは😂 いつもありがとうございます🙏 Strata v0.1.42 リリース 🚀 NVIDIA decode高速化、マルチGPU/AMD/Intelの安定化! ・NVIDIAのdecodeを約 +13〜16% 高速化 ・CPU/PCIe分担の自動最適化で条件次第 +30% ・RTX 5070 + Q2_0で 77.0 → 85.6 tok/s(+11.2%) ・マルチGPUの自動配置・layer splitを改善 ・3〜4GPU向けpipeline最適化でcode最大 +33.7%(opt-in) ・Data Parallel Replicaに対応 ・AMD gfx12 / R9700向けprompt高速化 ・Windows AMDの長文prompt時クラッシュ対策 ・AMD multi-GPUの104K-token checkpoint不具合を修正 ・Intel Arc B70/A750周りを多数修正・高速化…
Digest-only: model.name missing; incomplete methodology (context length, hardware state, measurement method).
RTX 3090 × 27B
Oct 10, 2026 · by @superalesha
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | — | — | 70.0 tok/s | — |
This post has 200K+ of views, and this is how local AI gets a bad name. "Run a 70B model on a 4GB GPU. 405B on 8GB." Yes, AirLLM can do that. It loads 1 layer from disk, computes it, drops it and loads the next. For every single token. Do the math. A 70B model is about 140GB of weights. To make 1 token you read all of it from disk. On a fast NVMe thats most of a minute per token. A 405B is about 800GB per token. Go…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, measurement method).
Unspecified platform × Qwen
Oct 10, 2026 · by @aitrackerbot · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 600K | — | — | 230.0 tok/s | — |
MISSINGNO. results: It speaks OpenAI's gpt-oss Harmony format natively, but on Qwen's tokenizer. Knows events through April 2026, takes 600K+ token prompts, ~230 tok/s. Matches none of the 880+ models we've fingerprinted. Not Kimi, DeepSeek, GLM or MiniMax. Maker still hidden.
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
DGX Spark x4 × GLM-5.3
Oct 10, 2026 · by @majewskizby
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 256K | INT4–INT8 Mix + NVFP4 sidecars | sparkdash | 40.0 tok/s | — |
| 256K | INT4–INT8 Mix + NVFP4 sidecars | sparkdash | 46.0 tok/s | — |
| 256K | INT4–INT8 Mix + NVFP4 sidecars | rigmark | 30.4 tok/s | — |
| 256K | INT4–INT8 Mix + NVFP4 sidecars | rigmark | 42.3 tok/s | — |
| 256K | INT4–INT8 Mix + NVFP4 sidecars | rigmark | 47.7 tok/s | — |
Full GLM-5.3 (753B) on 4x DGX Spark - update 🐁 • Quant: INT4–INT8 Mix + NVFP4 sidecars (attention/shared/dense/MTP); FP4x KV • Context: 256K • SparkDash c1: prose 35 → ~40 tok/s, code 42→ ~46 • RigMark c1 (thinking on): prose 28.4 → 30.4, code 39 → 42.3, structured 44.4 → 47.7 This is probably my last update for this setup. I fully expect others (@mmastrac, @MiaAI_lab) to crush these numbers soon - and I honestly…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
Also in this window
The 12 most recent posts are shown in full above; 15 posts from the same window are listed here in short form, each with its source post.
- RTX 5060 Ti x2 — 4176.0 tok/s prefill (1 measurement) · @llm_frontier_jp
- DGX Spark x4 × GLM-5.3-Flash — 286.7 tok/s decode (4 measurements) · @niklaslenz_ai
- DGX Spark x2 × GLM 5.3 Flash — 53.1 tok/s decode (2 measurements) · @atasrb
- Unspecified platform × DeepSeek V4 Pro — 1157.0 tok/s decode (1 measurement) · @OptionKing666
- Mac Mini M5 Pro × Qwen3 30B — 106.0 tok/s decode (6 measurements) · @pratikg
- RTX 3090 x4 × GLM-5.3 Flash — 79.0 tok/s decode (3 measurements) · @superalesha
- DGX Spark × Qwen3.8-Flash-Next — 75.7 tok/s decode (6 measurements) · @filicroval
- DGX Spark x4 × GLM-5.3 — 41.4 tok/s decode (3 measurements) · @niklaslenz_ai
- RTX 3060 — 89.9 tok/s decode (3 measurements) · @sudoingX
- DGX Spark x3, DGX Spark x4 × GLM5.3 Flash — 438.0 tok/s decode (1 measurement) · @losterror501
- RTX 5090 × Nemotron 3.5 Lightning — 731.0 tok/s decode (4 measurements) · @niklaslenz_ai
- DGX Spark x2 × DeepSeek-V4.1-Flash-EXL3 — 254.0 tok/s decode (4 measurements) · @sfxnz
- M3 Ultra × GLM-5.3-Flash — 40.6 tok/s decode (1 measurement) · @ivanfioravanti
- DGX Spark x4 × GLM 5.3 Flash — 222.0 tok/s decode (4 measurements) · @MiaAI_lab
- RTX 5090 × Qwen3.8 — 366.0 tok/s decode (3 measurements) · @MiaAI_lab
How to read this page
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and a number that turns out to be misreported is corrected or removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @ashxhart — Oct 11, 2026
- @ai_hakase_ — Oct 11, 2026
- @murasametech — Oct 11, 2026
- @wei_wang — Oct 11, 2026
- @ViC305 — Oct 11, 2026
- @murasametech — Oct 11, 2026
- @posi_posi8 — Oct 11, 2026
- @ScottLeimroth — Oct 11, 2026
- @jun1228909 — Oct 10, 2026
- @superalesha — Oct 10, 2026
- @aitrackerbot — Oct 10, 2026
- @majewskizby — Oct 10, 2026
- @llm_frontier_jp — Oct 10, 2026
- @niklaslenz_ai — Oct 10, 2026
- @atasrb — Oct 10, 2026
- @OptionKing666 — Oct 10, 2026
- @pratikg — Oct 10, 2026
- @superalesha — Oct 10, 2026
- @filicroval — Oct 10, 2026
- @niklaslenz_ai — Oct 10, 2026
- @sudoingX — Oct 10, 2026
- @losterror501 — Oct 10, 2026
- @niklaslenz_ai — Oct 10, 2026
- @sfxnz — Oct 10, 2026
- @ivanfioravanti — Oct 10, 2026
- @MiaAI_lab — Oct 10, 2026
- @MiaAI_lab — Oct 10, 2026
FAQ
Are the numbers in this digest sourced?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a row to reach the sourced tier?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Posts that state all five are published to the benchmark board as soon as the collector reads them, each row carrying its source link — nothing waits in a human review queue. Anything missing a condition stays in this digest.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.