DiggTop ⌕

X local-model bench digest — Oct 4, 2026

As of Oct 4, 2026 — 12 posts collected from X. 30 measured runs catalogued, 12 kept to the digest only. Latest source post Oct 4, 2026 10:40 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 16 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark × Kolibri-1

Oct 4, 2026 · by @0xBakeer

Context Quant Engine Decode Prefill
1K NVFP4 tandemllm 80.0 tok/s —
32K NVFP4 tandemllm 65.0 tok/s —
1K NVFP4 tandemllm — 2170.0 tok/s

If you still want to try Kolibri-1 on TandemLLM, I made an NVFP4 quant for it and tuned my engine around it. On one DGX Spark it decodes at 80 tok/s at 1k context and 65 at 32k, and prefill runs at about 2,170 tok/s. But I won't merge the branch. In my coding tests it thought only 200 to 300 tokens per turn, and the code came out poor. A drafter I trained for it also stayed too weak to pay off. For chat and German…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash

Oct 4, 2026 · by @NFT_Chen

Context Quant Engine Decode Prefill
262K EXL3 ~2.77 bpw vllm — 2000.0 tok/s
262K EXL3 ~2.77 bpw vllm 29.0 tok/s —
262K EXL3 ~2.77 bpw vllm 41.0 tok/s —

💥厉害!两台小盒子,跑起 552B 满血版 DeepSeek V4.1 Flash! DeepSeek-V4.1-Flash 不用再排队官方 API 了。只需用两台 DGX Spark + 一根 QSFP,把满血 Flash 拉到本地:EXL3 专家约 2.77 bpw,视觉、工具调用、reasoning 都在,还能直接接 Pi。 实测数据: 🔹Prefill ~2000 tok/s
🔹散文 decode ~29 tok/s,代码 ~41 tok/s
🔹KV 池约 200 万 token,单请求 262k
🔹Top-1 一致率约 92%,KLD 约 0.07
🔹175k 长上下文召回过关 想复现,按这个顺序: 1️⃣git clone https://t.co/FsZeeI47dS 2️⃣export WORKER=user@对端Spark的RoCE地址 3️⃣scripts/download.sh 拉权重并重组 Engram…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash

Oct 4, 2026 · by @bertholomusai

Context Quant Engine Decode Prefill
— — vllm v0.2 — 1357.0 tok/s
— — vllm v0.2 69.9 tok/s —
— — vllm v0.2 80.5 tok/s —

The first post left a lot on the table, so we went back to the formula. DeepSeek-V4.1-Flash on 2× DGX Spark, v0.2: • prefill 1,357→1,935 tok/s (+43%) • 1-stream code 69.9→80.4 (+15%) • 4 streams 80.5→~92 (+14%) • start 65→35 s Quality matters - bit for bit output. Hope you like speed, cuz we're not stopping here. https://t.co/YAN1K4SBAm

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, context length, hardware state, measurement method).

Unspecified platform × qwen3-1

Oct 4, 2026 · by @kayleecodez · third-party report

Context Quant Engine Decode Prefill
152K bf16 — 34.0 tok/s —

been optimizing the inference runtime im building (marin: rust + cuda, qwen3-1.7b bf16, batch 1). three measured fixes including my struggles: 1. sampling: for greedy decoding the cpu only needs the max logit but i was sorting all ~152k logits every token swapped it for one scan: 3.7ms - 0.15ms/token (~25x). decode barely moved tho. removing that cpu gap kept the gpu busy almost nonstop. then it hit the power cap…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, hardware state, measurement method).

M6 × Qwen3.8-27B

Oct 4, 2026 · by @anemll

Context Quant Engine Decode Prefill
65536 — core ai revision 1192a9c 196.0 tok/s —
65536 — core ai revision 1192a9c — 196.0 tok/s

ANEMLL Forge: Qwen3.8-27B on the M6 Neural Engine just got faster ! Full model, 8K–64K context: • Prefill 40–53% faster • speculative verify 24–45% faster • 64K-token prompt: 129 → 196 tok/s Quality holds (64K ppl 5.0686 → 5.0673) Update server https://t.co/1Y5F13miHW and model https://t.co/zcX7icXkQH ( Note, first compile takes ~22 minutes ) Main optimizations are DGN w/verifier: https://t.co/camVY6TyJZ

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (quantization, hardware state, measurement method).

M5 Ultra 256GB × DeepSeek-V4-Flash

Oct 4, 2026 · by @ashen_one

Context Quant Engine Decode Prefill
200K — — 37.6 tok/s —
200K — — — 658.2 tok/s

few more m5 ultra 256gb tests before i launch the full ashen benchmark 🫡 deepseek v4 flash on my m5 ultra still does 37.61 tok/s with 200k input 68% of its 8k speed with roughly 24x the input but processing that prompt takes 303.859 seconds before generation the wait before https://t.co/r6QiqNg4fT

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark x4 × GLM-5.3

Oct 4, 2026 · by @bertholomusai

Context Quant Engine Decode Prefill
1M EXL3 3.0 bpw tensorfold release build 82.0 tok/s —
1M EXL3 3.0 bpw tensorfold release build 64.0 tok/s —
1M EXL3 3.0 bpw tensorfold release build 40.0 tok/s —

Full GLM-5.3 (753B) on 4× DGX Spark. Our own 3.0 bpw EXL3 quant, served TP4 on TensorFold!! Measured on the release build: • 82 tok/s aggregate across 4 concurrent streams, TTFT ~0.5 s • ~64 tok/s aggregate with all 4 streams sharing the full 1M-token window • ~40 tok/s single stream • 1M context: needle found at 1,039,064 tokens • KL 0.111 vs BF16, top-1 agreement 90% • Every concurrent reply bit-identical to its…

Digest-only: incomplete methodology (hardware state, measurement method).

DGX Spark × Qwen3.8 Flash Next

Oct 4, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— — tensorfold 42.4 tok/s —
— — tensorfold 33.2 tok/s —
— — tensorfold 46.4 tok/s —
— — tensorfold 41.4 tok/s —

Qwen 3.8 Flash is starting to look very different depending on what inference stack you put underneath it. The model itself hasn’t changed. The interesting part is how much faster the same model can become when the runtime starts exploiting speculative decoding, context reuse and parallel verification. TensorFold is a good example. Its Qwen3.8 Flash Next setup uses an MTP head to draft several tokens ahead, then…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash

Oct 4, 2026 · by @bertholomusai

Context Quant Engine Decode Prefill
— EXL3 2.9bpw tensorfold release build 60.5 tok/s —
— EXL3 2.9bpw tensorfold release build 74.5 tok/s —
128K EXL3 2.9bpw tensorfold release build 61.8 tok/s 1350.0 tok/s
— EXL3 2.9bpw tensorfold release build 80.0 tok/s —

DeepSeek-V4.1-Flash on just 2× DGX Spark — with our own engine. We wrote a clean-room deepseek_v41 family for TensorFold (by @ashxhart): tensor-parallel over two GB10s, from DeepSeek's MIT inference code + tech report. Measured on the release build: • 60.5 tok/s code / 74.5 structured, single stream (DSpark) • ~80 tok/s across 4 concurrent streams — every reply bit-identical to its solo run • ~1,350 tok/s prefill…

Digest-only: incomplete methodology (hardware state, measurement method).

M5 Max × DeepSeek V4 Flash

Oct 3, 2026 · by @RoundtableSpace · third-party report

Context Quant Engine Decode Prefill
— — — 39.0 tok/s —

Local AI tip: if you've got a big-memory Mac, you can run frontier open models at home now. antirez, the creator of Redis, built ds4, a lean C engine that runs a few top open models on your own machine. the setup: * models → DeepSeek V4 Flash, GLM 5.x and Qwen 3.8 Flash Next * hardware → high-memory Macs, NVIDIA and AMD GPUs * speed → about 39 tokens per second on a 128GB M5 Max * tools → a CLI, an OpenAI and…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x1 × Kolibri 1

Oct 3, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
164K FP8 vllm 0.29.0 48.0 tok/s —
164K FP8 vllm 0.29.0 157.0 tok/s —
164K FP8 vllm 0.29.0 — 5200.0 tok/s

Got Aleph Alpha's new Kolibri 1 running on a single DGX Spark the day it came out. Their container is x86 only, so I built vLLM 0.29.0 (arm64) + their aleph-alpha-inference plugin. FP8 weights (74 GB) and an FP8 KV cache, 164K context. The FP8 kernels worked on GB10 even though torch doesn't list sm_121. Speed: 48 tok/s solo, 157 tok/s with 8 streams, 5.2K tok/s prompt read. Then I gave it the pagoda prompt 👇…

Digest-only: incomplete methodology (measurement method).

Unspecified platform

Oct 3, 2026 · by @Da7_Tech · third-party report

Context Quant Engine Decode Prefill
84K — — 130.0 tok/s —

SWE-2 after three weeks of real work. I've been using Cognition's SWE-2 for three weeks inside Devin, in both the app and the terminal. I talked to it directly, ran it as a subagent under Opus 5.5 and Astra, and let it orchestrate its own copies. Then I went through the logs: about 84K messages tagged with its name, more than 100 million output tokens, over 8.4 billion tokens in total, and more than a thousand runs…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; model.name missing; incomplete methodology (engine+version, quantization, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑