X local-model bench digest — Oct 4, 2026
As of Oct 4, 2026 — 12 posts collected from X. 30 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 4, 2026 10:40 UTC.
Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.
The 12 most recent posts are listed below; 16 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).
DGX Spark × Kolibri-1
Oct 4, 2026 · by @0xBakeer
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1K | NVFP4 | tandemllm | 80.0 tok/s | — |
| 32K | NVFP4 | tandemllm | 65.0 tok/s | — |
| 1K | NVFP4 | tandemllm | — | 2170.0 tok/s |
If you still want to try Kolibri-1 on TandemLLM, I made an NVFP4 quant for it and tuned my engine around it. On one DGX Spark it decodes at 80 tok/s at 1k context and 65 at 32k, and prefill runs at about 2,170 tok/s. But I won't merge the branch. In my coding tests it thought only 200 to 300 tokens per turn, and the code came out poor. A drafter I trained for it also stayed too weak to pay off. For chat and German…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).
DGX Spark x2 × DeepSeek-V4.1-Flash
Oct 4, 2026 · by @NFT_Chen
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 262K | EXL3 ~2.77 bpw | vllm | — | 2000.0 tok/s |
| 262K | EXL3 ~2.77 bpw | vllm | 29.0 tok/s | — |
| 262K | EXL3 ~2.77 bpw | vllm | 41.0 tok/s | — |
💥厉害!两台小盒子,跑起 552B 满血版 DeepSeek V4.1 Flash! DeepSeek-V4.1-Flash 不用再排队官方 API 了。只需用两台 DGX Spark + 一根 QSFP,把满血 Flash 拉到本地:EXL3 专家约 2.77 bpw,视觉、工具调用、reasoning 都在,还能直接接 Pi。 实测数据: 🔹Prefill ~2000 tok/s 🔹散文 decode ~29 tok/s,代码 ~41 tok/s 🔹KV 池约 200 万 token,单请求 262k 🔹Top-1 一致率约 92%,KLD 约 0.07 🔹175k 长上下文召回过关 想复现,按这个顺序: 1️⃣git clone https://t.co/FsZeeI47dS 2️⃣export WORKER=user@对端Spark的RoCE地址 3️⃣scripts/download.sh 拉权重并重组 Engram…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).
DGX Spark x2 × DeepSeek-V4.1-Flash
Oct 4, 2026 · by @bertholomusai
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | vllm v0.2 | — | 1357.0 tok/s |
| — | — | vllm v0.2 | 69.9 tok/s | — |
| — | — | vllm v0.2 | 80.5 tok/s | — |
The first post left a lot on the table, so we went back to the formula. DeepSeek-V4.1-Flash on 2× DGX Spark, v0.2: • prefill 1,357→1,935 tok/s (+43%) • 1-stream code 69.9→80.4 (+15%) • 4 streams 80.5→~92 (+14%) • start 65→35 s Quality matters - bit for bit output. Hope you like speed, cuz we're not stopping here. https://t.co/YAN1K4SBAm
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).
Unspecified platform × qwen3-1
Oct 4, 2026 · by @kayleecodez · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 152K | bf16 | — | 34.0 tok/s | — |
been optimizing the inference runtime im building (marin: rust + cuda, qwen3-1.7b bf16, batch 1). three measured fixes including my struggles: 1. sampling: for greedy decoding the cpu only needs the max logit but i was sorting all ~152k logits every token swapped it for one scan: 3.7ms - 0.15ms/token (~25x). decode barely moved tho. removing that cpu gap kept the gpu busy almost nonstop. then it hit the power cap…
Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, hardware state, measurement method).
M6 × Qwen3.8-27B
Oct 4, 2026 · by @anemll
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 65536 | — | core ai revision 1192a9c | 196.0 tok/s | — |
| 65536 | — | core ai revision 1192a9c | — | 196.0 tok/s |
ANEMLL Forge: Qwen3.8-27B on the M6 Neural Engine just got faster ! Full model, 8K–64K context: • Prefill 40–53% faster • speculative verify 24–45% faster • 64K-token prompt: 129 → 196 tok/s Quality holds (64K ppl 5.0686 → 5.0673) Update server https://t.co/1Y5F13miHW and model https://t.co/zcX7icXkQH ( Note, first compile takes ~22 minutes ) Main optimizations are DGN w/verifier: https://t.co/camVY6TyJZ
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, hardware state, measurement method).
M5 Ultra 256GB × DeepSeek-V4-Flash
Oct 4, 2026 · by @ashen_one
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 200K | — | — | 37.6 tok/s | — |
| 200K | — | — | — | 658.2 tok/s |
few more m5 ultra 256gb tests before i launch the full ashen benchmark 🫡 deepseek v4 flash on my m5 ultra still does 37.61 tok/s with 200k input 68% of its 8k speed with roughly 24x the input but processing that prompt takes 303.859 seconds before generation the wait before https://t.co/r6QiqNg4fT
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).
DGX Spark x4 × GLM-5.3
Oct 4, 2026 · by @bertholomusai
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 1M | EXL3 3.0 bpw | tensorfold release build | 82.0 tok/s | — |
| 1M | EXL3 3.0 bpw | tensorfold release build | 64.0 tok/s | — |
| 1M | EXL3 3.0 bpw | tensorfold release build | 40.0 tok/s | — |
Full GLM-5.3 (753B) on 4× DGX Spark. Our own 3.0 bpw EXL3 quant, served TP4 on TensorFold!! Measured on the release build: • 82 tok/s aggregate across 4 concurrent streams, TTFT ~0.5 s • ~64 tok/s aggregate with all 4 streams sharing the full 1M-token window • ~40 tok/s single stream • 1M context: needle found at 1,039,064 tokens • KL 0.111 vs BF16, top-1 agreement 90% • Every concurrent reply bit-identical to its…
Digest-only: incomplete methodology (hardware state, measurement method).
DGX Spark × Qwen3.8 Flash Next
Oct 4, 2026 · by @Oluwaphilemon1
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | tensorfold | 42.4 tok/s | — |
| — | — | tensorfold | 33.2 tok/s | — |
| — | — | tensorfold | 46.4 tok/s | — |
| — | — | tensorfold | 41.4 tok/s | — |
Qwen 3.8 Flash is starting to look very different depending on what inference stack you put underneath it. The model itself hasn’t changed. The interesting part is how much faster the same model can become when the runtime starts exploiting speculative decoding, context reuse and parallel verification. TensorFold is a good example. Its Qwen3.8 Flash Next setup uses an MTP head to draft several tokens ahead, then…
Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x2 × DeepSeek-V4.1-Flash
Oct 4, 2026 · by @bertholomusai
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | EXL3 2.9bpw | tensorfold release build | 60.5 tok/s | — |
| — | EXL3 2.9bpw | tensorfold release build | 74.5 tok/s | — |
| 128K | EXL3 2.9bpw | tensorfold release build | 61.8 tok/s | 1350.0 tok/s |
| — | EXL3 2.9bpw | tensorfold release build | 80.0 tok/s | — |
DeepSeek-V4.1-Flash on just 2× DGX Spark — with our own engine. We wrote a clean-room deepseek_v41 family for TensorFold (by @ashxhart): tensor-parallel over two GB10s, from DeepSeek's MIT inference code + tech report. Measured on the release build: • 60.5 tok/s code / 74.5 structured, single stream (DSpark) • ~80 tok/s across 4 concurrent streams — every reply bit-identical to its solo run • ~1,350 tok/s prefill…
Digest-only: incomplete methodology (hardware state, measurement method).
M5 Max × DeepSeek V4 Flash
Oct 3, 2026 · by @RoundtableSpace · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| — | — | — | 39.0 tok/s | — |
Local AI tip: if you've got a big-memory Mac, you can run frontier open models at home now. antirez, the creator of Redis, built ds4, a lean C engine that runs a few top open models on your own machine. the setup: * models → DeepSeek V4 Flash, GLM 5.x and Qwen 3.8 Flash Next * hardware → high-memory Macs, NVIDIA and AMD GPUs * speed → about 39 tokens per second on a 128GB M5 Max * tools → a CLI, an OpenAI and…
Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).
DGX Spark x1 × Kolibri 1
Oct 3, 2026 · by @WescheNex1q
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 164K | FP8 | vllm 0.29.0 | 48.0 tok/s | — |
| 164K | FP8 | vllm 0.29.0 | 157.0 tok/s | — |
| 164K | FP8 | vllm 0.29.0 | — | 5200.0 tok/s |
Got Aleph Alpha's new Kolibri 1 running on a single DGX Spark the day it came out. Their container is x86 only, so I built vLLM 0.29.0 (arm64) + their aleph-alpha-inference plugin. FP8 weights (74 GB) and an FP8 KV cache, 164K context. The FP8 kernels worked on GB10 even though torch doesn't list sm_121. Speed: 48 tok/s solo, 157 tok/s with 8 streams, 5.2K tok/s prompt read. Then I gave it the pagoda prompt 👇…
Digest-only: incomplete methodology (measurement method).
Unspecified platform
Oct 3, 2026 · by @Da7_Tech · third-party report
| Context | Quant | Engine | Decode | Prefill |
|---|---|---|---|---|
| 84K | — | — | 130.0 tok/s | — |
SWE-2 after three weeks of real work. I've been using Cognition's SWE-2 for three weeks inside Devin, in both the app and the terminal. I talked to it directly, ran it as a subagent under Opus 5.5 and Astra, and let it orchestrate its own copies. Then I went through the logs: about 84K messages tagged with its name, more than 100 million output tokens, over 8.4 billion tokens in total, and more than a thousand runs…
Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).
Verification status
This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.
Sources
- @0xBakeer — Oct 4, 2026
- @NFT_Chen — Oct 4, 2026
- @bertholomusai — Oct 4, 2026
- @kayleecodez — Oct 4, 2026
- @anemll — Oct 4, 2026
- @ashen_one — Oct 4, 2026
- @bertholomusai — Oct 4, 2026
- @Oluwaphilemon1 — Oct 4, 2026
- @bertholomusai — Oct 4, 2026
- @RoundtableSpace — Oct 3, 2026
- @WescheNex1q — Oct 3, 2026
- @Da7_Tech — Oct 3, 2026
FAQ
Are the numbers in this digest verified?
They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.
What does it take for a run to reach the verified board?
Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.
Where does the data come from?
Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.