DiggTop ⌕

X local-model bench digest — Oct 5, 2026

As of Oct 5, 2026 — 12 posts collected from X. 32 measured runs catalogued, 1 queued for review, 11 kept to the digest only. Latest source post Oct 5, 2026 13:27 UTC.

Every number below is what its author published, transcribed with the engine, quantization and context length they reported. Sources are linked on each entry. Numbers without a full methodology stay out of the benchmark board until they are reviewed.

The 12 most recent posts are listed below; 13 older entries from this window are in the daily collection report (a digest-only entry never reaches the database).

DGX Spark x2 × DeepSeek-V4.1-Flash

Oct 5, 2026 · by @jayleaton

Context Quant Engine Decode Prefill
— — tensorfold — 2310.0 tok/s
32768 — tensorfold — 2025.0 tok/s
— — tensorfold 85.0 tok/s —
— — tensorfold 47.0 tok/s —
— — tensorfold 122.0 tok/s —
— — tensorfold 100.0 tok/s —

DeepSeek 4.1 Flash TensorFold 2x DGX Spark update: In the last update, I said I was chasing a memory leak and hoped for another 10-15% prompt speed. Well, I got it. Replay mode: +13%. (2,310) Exact mode doubled! 32K prompt: 1,004 → 2,025 tok/s Same output, token for token. Everything else moved a bit too: code 82.8 → ~85 tok/s prose 44.3 → ~47 structured 118 → ~122 4 streams 96.6 → ~100 And the leak: 1 hour soak…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash

Oct 5, 2026 · by @sudoingX

Context Quant Engine Decode Prefill
— EXL3 2.77bpw vllm fork with exllamav3 and b12x kernels 34.8 tok/s —
— EXL3 2.77bpw vllm fork with exllamav3 and b12x kernels 22.5 tok/s —

i always wanted deepseek running on my sparks and tonight it's loaded, @0xSero's v4.1 flash recipe on my 2x dgx sparks i built the serving image from source instead of pulling it, the vllm fork with exllamav3 and the b12x kernels, pulled the 462gb checkpoint onto both boxes and it came up over the connectx cable in about 15 minutes the checkpoint is wild, the routed experts squeezed to 2.77 bits with exl3 and the…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (context length, hardware state, measurement method).

DGX Spark x2 × GLM-5.3-Flash

Oct 5, 2026 · by @sudoingX

i had two of the best local models build the same floating tree, glm 5.3 flash at 320b and qwen 3.8 flash next at 125b, and despite the size gap look where both landed glm 5.3 flash: 320b moe with 18b active, nvidia's nvfp4 weights split over 2x dgx spark, served with vllm and mtp k=3 at 20 to 30 tok/s, 131k context, hermes agent driving. 12/12 tests at 1h34m, done at 4.5 hours with 2,356 lines across 8 files…

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

M3 Ultra

Oct 5, 2026 · by @qvac

Context Quant Engine Decode Prefill
— — llama.cpp 110.0 tok/s —
— — llama.cpp 32.1 tok/s —

One of our PR just got merged into llama.cpp: new Metal kernels for speculative decoding on Apple Silicon. Before this PR, speculative decoding on a Mac was slower than plain decoding. On an M3 Ultra it now runs up to 3.4x faster than plain decoding, 110 tok/s against 32.1. Thanks to @ggerganov for reviewing and refining it. https://t.co/Q0Zlcf9K0G

Digest-only: model.name missing; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 5090 × Qwen3.8-27B

Oct 5, 2026 · by @Oluwaphilemon1 · third-party report

Context Quant Engine Decode Prefill
— — — 119.0 tok/s —
— — — 103.0 tok/s —

Mitsuba-ComfyUI just got a very interesting add-on. HiMitsuba is a 70MB LoRA for the 7.3GB PQ2_0 Mitsuba model, which itself is a ternary 1.58-bit version of Qwen3.8-27B tuned specifically for ComfyUI and vision-to-prompt workflows. (Hugging Face) The clever part isn’t the size. It’s what the LoRA changes. Mitsuba is designed for a very specific job: Show it an image and have it create a usable image or video…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, context length, hardware state, measurement method).

RTX 6000 PRO x2 × GLM 5.3 Flash

Oct 5, 2026 · by @MiaAI_lab

Context Quant Engine Decode Prefill
512K EXL3 tensorfold 238.0 tok/s —
512K EXL3 tensorfold 449.0 tok/s —

GLM 5.3 Flash EXL3 for 2x RTX 6000 PROs 🔥 - 238 tok/s single stream prose - 449 tok/s at 4 stream prose - 1M KV pool and 512K context Check it out! https://mia-ai.net/models/GLM-5.3-Flash-EXL3-2x-RTX-PRO-6000-TensorFold

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x4 × GLM-5.3 TP4

Oct 5, 2026 · by @WescheNex1q

Context Quant Engine Decode Prefill
— — tensorfold 40.7 tok/s —
— — tensorfold 43.4 tok/s —
— — tensorfold 36.6 tok/s —
— — tensorfold 40.6 tok/s —
— — tensorfold 27.1 tok/s —
— — tensorfold 35.4 tok/s —

Reproduced @u1tra_instinct new full GLM-5.3 TP4 image with Tensorfold on our own 4× DGX Spark Two runs: prose: 40.7 / 43.4 tok/s (his 41.8) code: 36.6 / 40.6 tok/s (his 38.1) 25K-token prefill + needle: 22.7–24.0 s (his ~25 s), PASS clocks 2.4–2.5 GHz, 68–71 °C, 0 swap Long reasoning: 27.1 → 35.4 tok/s vs the previous image.

Digest-only: no run with engine+version+quant+context+decode tps; incomplete methodology (engine+version, quantization, context length, measurement method).

RTX 3090 × Qwen

Oct 5, 2026 · by @jurlycat · third-party report

Context Quant Engine Decode Prefill
300K — — 60.0 tok/s —
300K — — 30.0 tok/s —
300K — — 20.0 tok/s —

A Reddit user started with one RTX 3090, then eventually scaled to 16 RTX 3090s. That rig reportedly ran Qwen 397B at 50–60 tok/s while drawing ~6 kW. It ran fine… until someone turned on the oven. Running both overloaded the house circuit and blew the fuses. 💀 He then switched to 4 ASUS GB10s: ~30 tok/s on the same model at ~400 W. Later, he and his brother combined two 8-machine ASUS GB10 clusters. Their…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, quantization, hardware state, measurement method).

DGX Spark × GLM-5.3

Oct 5, 2026 · by @WescheNex1q · third-party report

Context Quant Engine Decode Prefill
125600 EXL3 — 30.0 tok/s —

GLM 753B vs Flash, same voxel pagoda prompt, both EXL3 on TensorFold, all local -GLM-5.3 full (4× DGX Spark): 125.6K tokens, 72 min, ran first try -GLM-5.3 Flash (2× DGX Spark): 103.4K tokens, 58 min, 1 fix Same speed, about 30 tok/s. The big one got it right first time; Flash brought the color. Setup: full GLM-5.3 = @u1tra_instinct ' EXL3 2.75bpw + TensorFold TP4 recipe. Flash = @MiaAI_lab EXL3 TR3-4bpw. Both on…

Digest-only: confidence 0.35 below 0.60; incomplete methodology (engine+version, hardware state, measurement method).

DGX Spark x2 × DeepSeek-V4.1-Flash-EXL3-2.9bpw

Oct 5, 2026 · by @bertholomusai

Context Quant Engine Decode Prefill
384 EXL3 2.9bpw tensorfold bd0024d 101.1 tok/s —
384 EXL3 2.9bpw tensorfold bd0024d 62.5 tok/s —
384 EXL3 2.9bpw tensorfold bd0024d 142.4 tok/s —
128K EXL3 2.9bpw tensorfold bd0024d 89.2 tok/s —
384 EXL3 2.9bpw tensorfold bd0024d 124.5 tok/s —
128K EXL3 2.9bpw tensorfold bd0024d — 1732.0 tok/s

@deepseek_ai V4.1-Flash on 2x DGX Spark, v0.4: single stream past 100 tok/s. • code 86.0 → 101.1 tok/s (+18%) • prose 53.2 → 62.5 (+18%) • 4 streams sustained 107.5 → 124.5 (+16%) • decode at 128K 59.9 → 89.2 (+49%) Bit-identical to v0.3. https://github.com/bertholomus/deepseek-v4.1-tensorfold-tp2-2xgb10 https://t.co/1FQOGAZsN7

Full methodology stated (engine, quantization, context, hardware state, method) — 2 added to the benchmark board (2 of 6 measured rows queued; the rest are published here only).

Unspecified platform × Qwen3.8-Flash-Next

Oct 5, 2026 · by @Oluwaphilemon1 · third-party report

Context Quant Engine Decode Prefill
230K BF16 — 35.0 tok/s —

Qwen3.8-Flash-Next is doing something that looks almost wrong at first glance. You can have roughly 102GB of BF16 n-gram data sitting on an NVMe SSD, push the context out toward 230K tokens, and still get around 35 tokens/sec on a single Radeon setup. That sounds impossible if you think of the SSD as holding a normal 102GB chunk of model weights. But the n-gram table isn’t being used like a conventional neural…

Digest-only: confidence 0.35 below 0.60; hardware.name missing; incomplete methodology (engine+version, hardware state, measurement method).

Unspecified platform × Qwen3.8-Flash-Next

Oct 5, 2026 · by @Oluwaphilemon1

Context Quant Engine Decode Prefill
— IQ3_XXS strata 236.0 tok/s —

Qwen3.8-Flash-Next just got another interesting local inference setup. An abliterated build running through Strata with the Swift IQ3_XXS quantization successfully reached around 236 tok/s. That number is the headline. But the stack behind it is what makes the result interesting. You’re combining: Qwen3.8-Flash-Next Abliterated weights Swift fine-tune IQ3_XXS quantization Strata inference And getting extremely high…

Digest-only: hardware.name missing; incomplete methodology (engine+version, context length, hardware state, measurement method).

Verification status

This digest is an automated transcription of public posts. Runs with a stated engine build, quantization, context length, hardware state and measurement method are published straight to the benchmark board with a link to their source; the rest stay here. These are author-reported numbers, not an independent DiggTop re-measurement: every row carries its provenance so you can check the original, and numbers that fail review are removed. Our own testing follows the methodology. Spotted an error? Reply on the source post or submit a corrected run at /submit.

Sources

FAQ

Are the numbers in this digest verified?

They are transcribed from public X posts and the repositories those posts link to, with attribution and a source link on every entry. Rows that state the full methodology are published to the benchmark board automatically, so treat them as author-reported until a dispute is raised: the source link on every row is the check.

What does it take for a run to reach the verified board?

Engine and exact version, quantization, context length, hardware state and a described measurement method. Author-reported posts that state all five are queued as pending and can be promoted by a reviewer; anything missing a condition stays out of the database.

Where does the data come from?

Public X posts about local-model inference measured on owned hardware, discovered by keyword search every two hours and linked back to the original thread. The digest is regenerated daily.

↑