DiggTop ⌕

Self-measured runs

As of 2026-10-07. Of the 468 rows on the benchmark board, two were measured by us, on hardware we own. The other 466 are attributed claims: 361 imported from localmaxxing and 105 from public X posts, each linking back to its author. What a row means explains the three tiers.

Both of our runs are one machine: an RTX 4090 D 24GB workstation running Qwen3.8-27B under llama.cpp, single stream, warm server.

# Date Quant Engine Context Decode Prefill
1 2026-09-04 IQ3_S (GSQ-RCO) llama.cpp c56093 (AtomicBot turboquant fork) 262,144 67.4 tok/s —
2 2026-09-08 Q4_K_M llama.cpp b358-7620399 2,421 70.9 tok/s 1,467 tok/s

Run 1 — recorded parameters

Run 2 — what we did not record

Run 2 has the engine commit, the quantization, the context length and both throughput figures, and nothing else: no driver version, no batch settings, no speculative-decoding settings.

Neither run has a saved command line. We did not keep one, so we are not going to reconstruct one after the fact — a plausible-looking command in a table is worse than a known gap. Treat both rows as measured once, not independently repeated: they carry the same "no second source" caveat as every other row on the board (see methodology).

If you want to reproduce either, the honest starting point is: llama-bench (or llama-server) on a 4090 D, Qwen3.8-27B at the quant above, the stated context, and for run 1 the MTP head with n-max=2. Your numbers will differ with the driver, the batch settings and the exact build.

Why this page exists

Because "we verified it" is a claim we cannot support for other people's measurements, and it would be dishonest to make an exception for our own. Publishing the two rows we did run — with the gaps visible — is the only version of that claim we can stand behind.

Sources

↑