DiggTop

X local-model bench digest — Sep 19, 2026

As of Sep 19, 2026 — 12 posts collected from X. 49 measured runs catalogued, 12 kept to the digest only. Latest source post Sep 19, 2026 13:36 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 13 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x4 × DeepSeek-V4.1-Flash

Sep 19, 2026 · by @Tech2Wild

Context Quant Engine Decode Prefill
500K EXL3 3.5bpw 91.0 tok/s
500K EXL3 3.5bpw 74.0 tok/s
500K EXL3 3.5bpw 1378.0 tok/s
500K EXL3 3.5bpw 1401.0 tok/s
500K EXL3 3.5bpw 324.0 tok/s
500K EXL3 3.5bpw 88.0 tok/s
500K EXL3 3.5bpw 281.0 tok/s
500K EXL3 3.5bpw 75.0 tok/s
500K EXL3 3.5bpw 55.0 tok/s
500K EXL3 3.5bpw 70.0 tok/s
500K EXL3 3.5bpw 39.0 tok/s
500K EXL3 3.5bpw 33.0 tok/s

HUGE SPEED RUN ⚡ Performance improvements, DeepSeek V4.1 Flash on 4x DGX Spark (TP=4, EXL3 3.5bpw, 500K ctx): 💻 Code ~8% faster on 2-4 concurrent streams ~91 tok/s single stream · ~74 tok/s per stream at 2-4 · 324 tok/s aggregate at 6 🚀 🧮 Math ~9% faster on 2-4 concurrent streams ~88 tok/s single stream · 281 tok/s aggregate at 6 🧠 Reasoning ~13% faster single stream, ~11% at 2-4 streams ~75 tok/s single stream 📋…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

RTX 3060 × Qwen3.8-27B

Sep 19, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
262K llama.cpp prismml fork 26.0 tok/s
12K llama.cpp prismml fork 22.0 tok/s
35K llama.cpp prismml fork 18.0 tok/s
77K llama.cpp prismml fork 13.0 tok/s

the timeline spent two days saying bonsai 2 cannot build. here is 1 hour 24 minutes of it building, one shot, from one paragraph, on an rtx 3060 12gb, sped to 8x so you can watch the whole thing. what you are watching is a 5.9gb ternary compression of qwen 3.8 27b, served through the prismml llama.cpp fork with the full 262k window resident, writing a terminal gpu monitor in one html file. it thinks, it calls a tool…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

DGX Spark × Qwen3.8-27B

Sep 19, 2026 · by @eta1ia

Context Quant Engine Decode Prefill
NVFP4 vllm 15.9 tok/s
NVFP4 sglang 28.3 tok/s
NVFP4 vllm 24.4 tok/s
NVFP4 sglang 55.2 tok/s

DGX Spark で動かしている vLLM + Qwen3.8-27B NVFP4 + MTP の構成を SGLang + Qwen3.8-27B NVFP4 + DFlash2 に変更したらからり速度が改善した。 ・通常文章: 15.95 tok/s → 28.30 tok/s ・コーディング: 24.44 tok/s → 55.25 tok/s 個人で使うだけだしこれが良さそう。

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Tesla V100 PCIe 32GB × Qwen3.8-27B

Sep 19, 2026 · by @UtaAoya

Context Quant Engine Decode Prefill
NVFP4 ninfer v100 29.0 tok/s
NVFP4 dflash2 54.0 tok/s

Local LLM - V100環境-150W縛り 自作PCで、まずは素振りベンチマーク。 V100 x 1 (PG500-216 / GV100GL) なので、単体ではNInferが最速らしい?(9/19 チャッピー調べ) コンテキストを伸ばしてのテストはこれから😊 チャッピー解説 Tesla V100 PCIe 32GB(PG500-216 / GV100GL)単体を、150W縛りで速度比較。 Qwen3.8-27B NVFP4 + ninfer-v100で、通常生成(no-spec)の 約29 tok/s に対し、DFlash2では 最大約54 tok/s を確認。 V100を150Wに制限した状態でも約1.86倍まで高速化できたのがポイント。一方、投機デコードは設定によって出力一致性に差が出るため、そこは継続検証中です。 #地元のV100クラブ

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, measurement method).

iPhone 18 Pro Max × Gemma 4 12B

Sep 19, 2026 · by @bijanbowen

Context Quant Engine Decode Prefill
Q4_0 14.0 tok/s
Q4_K_M 57.7 tok/s
Q8_0 135.5 tok/s

iPhone 18 Pro Max Local LLM: Gemma 4 12B (Q4_0): ~14 tok/s Granite 4.0 H Tiny (Q4_K_M): 57.7 tok/s Qwen 3 0.6B (Q8_0): 135.5 tok/s https://t.co/ff7yVxMKqX

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

Mac Studio M4 Max 128GB × Qwen3.8-27B

Sep 19, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
1024 4-bit splash 68.0 tok/s
1024 4-bit mlx 48.9 tok/s
1024 4-bit splash 53.7 tok/s
1024 4-bit mlx 34.7 tok/s
1024 4-bit splash 33.9 tok/s
1024 4-bit mlx 28.4 tok/s
1024 4-bit splash 36.3 tok/s
1024 4-bit mlx 38.3 tok/s

Splash vs mlx-vlm+MTP, Qwen3.8-27B 4-bit Mac Studio M4 Max 128 GB. Stock splash serve, nothing else on the GPU, 5 coding prompts, 1,024-token cap, medians: Code, no reasoning: Splash 68.0 · MLX 48.9 Code, reasoning on: 53.7 · 34.7 32K+ thinking run: 33.9 · 28.4 Short prose: 36.3 · 38.3 Splash is 20–40% faster on code on this box. Ties on prose. Their 74 tok/s is an M5 Pro number; M4 Max lands 54–68 depending on…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).

RTX 3060 × Ternary Bonsai 2 27B

Sep 19, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
7K llama.cpp 24.4 tok/s
12K llama.cpp 21.9 tok/s
35K llama.cpp 17.8 tok/s
77K llama.cpp 13.0 tok/s
41312 llama.cpp 15.0 tok/s
2K llama.cpp 295.0 tok/s
35K llama.cpp 243.0 tok/s
llama.cpp 26.1 tok/s

Bonsai 2 27B on an RTX 3060 12GB. Full receipt. This might be one of the more useful local AI tests for anyone sitting on a 12 GB GPU. The model is Ternary Bonsai 2 27B, a compressed version of Qwen3.8-27B. And yes, the entire native 262K context window fits on a 12 GB RTX 3060. Here are the numbers from a live server with thinking enabled: Generation speed by context depth 7K: 24.4 tok/s 12K: 21.9 tok/s 35K: 17.8…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version).

DGX Spark × GLM-5.3

Sep 18, 2026 · by @mr_r0b0t · third-party report

Context Quant Engine Decode Prefill
EXL3 58.4 tok/s

Cooked up a GLM-5.3-Flash-EXL3–2.32bpw for the 128GB crowd, perfect for your @NVIDIAAI DGX Spark ⚡️ Testing it now, looks like quantized DFlash2-EXL3 is working well! 58.4 tok/s isn’t to be taken as any sort of eval here but it is a very positive signal 🤩 https://t.co/7JFtjPQuD3

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, context length, hardware state, measurement method).

A100 × deepseekv4.1-A100-custom

Sep 18, 2026 · by @shi3z

Context Quant Engine Decode Prefill
litellm 120.0 tok/s
1M litellm 50.0 tok/s
litellm 80.0 tok/s
litellm 12979.8 tok/s

Just updated the README for deepseekv4.1-A100-custom! 🚀 Added benchmark specs & performance numbers, along with a full configuration and setup guide (including running on multi-A100 setups). Speed&Multi-Agent: 120tok/s 1M Context Robust 50tok/s Balanced 80tok/s jev_mode: 12,979.8 tok/s Change Claude Code backend from cloud to local via liteLLM Check it out here: 👉 https://t.co/P9RxERX0Ut

Digest-only: confidence 0.40 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

RTX 4090 × Bonsai 2 27B

Sep 18, 2026 · by @studio_yebisu

Context Quant Engine Decode Prefill
64.6 tok/s

昨日、Xのゆるトークとかあったから試せなかった Qwen3.8-27Bを3値量子化を施したマルチモーダルLLMであるBonsai 2 27Bを早速弊環境で試してみたんだが。 完全にローカルなのに、64.6 tok/s出るもんだから Asset無し(画像とか渡さず、Net接続もさせず)でなんかそれなりのWebページを25秒ぐらいで作りやがったんですね。 VRAMめっちゃ空いてるし。(12GBあれば十分すぎ) 本家のKimi K3が30tok/sとかだから、用途次第ではめっちゃ活用出来るじゃんって感じなんだよな。なんだよこの謎の3値量子化とかいう新技術。聞いたことねぇよ。 まぁ大分久しぶりの #地元のLLM だったもんだから、進化に驚いた、という話。 環境:7950X3D,RTX4090 24GB,DDR5 64GB

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, context length, hardware state, measurement method).

M4 Max

Sep 18, 2026 · by @rapidmlx · third-party report

Context Quant Engine Decode Prefill
mlx v0.14.3 34.0 tok/s

🚀 Rapid-MLX v0.14.3 is live! The star: Bonsai 2 Hadamard 27B. A multimodal 27B model packed into just ~8GB, blazing at 34 tok/s on an M4 Max. Full support across CLI, Server, and Desktop for vision, reasoning, and tool calling. Try it: rapid-mlx serve bonsai2-27b-2bit Huge shoutout to @Mieluoxxx (Morgan Woods) for the original Bonsai 2 loader contribution! 🙌 On top of that ⚡️ Multimodal conversations are now much…

Digest-only: confidence 0.35 below 0.60; model.name missing; incomplete methodology (quantization, context length, hardware state, measurement method).

DGX Spark × glm-5.3

Sep 18, 2026 · by @0xSero · third-party report

Context Quant Engine Decode Prefill
262K EXL3 20.0 tok/s

DGX Spark owners rejoice! Finally managed to make a better quant for glm-5.3-flash, you get 262k context, vision, and 10-20 tok/s on 1x DGX Spark. 80% top 1 agreement, lower KL, lower perplexity. Enjoy! Great as an overnight agent https://huggingface.co/0xSero/GLM-5.3-Flash-EXL3-Spark https://t.co/HmuRMzVroM

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.