Self-measured runs
As of 2026-10-07. Of the 468 rows on the benchmark board, two were measured by us, on hardware we own. The other 466 are attributed claims: 361 imported from localmaxxing and 105 from public X posts, each linking back to its author. What a row means explains the three tiers.
Both of our runs are one machine: an RTX 4090 D 24GB workstation running Qwen3.8-27B under llama.cpp, single stream, warm server.
| # | Date | Quant | Engine | Context | Decode | Prefill |
|---|---|---|---|---|---|---|
| 1 | 2026-09-04 | IQ3_S (GSQ-RCO) | llama.cpp c56093 (AtomicBot turboquant fork) |
262,144 | 67.4 tok/s | — |
| 2 | 2026-09-08 | Q4_K_M | llama.cpp b358-7620399 |
2,421 | 70.9 tok/s | 1,467 tok/s |
Run 1 — recorded parameters
- MTP speculative decoding n-max=2,
ubatch1024, flash attention on - Ubuntu 24.04, NVIDIA driver 580.173.02
- Warm server, single stream, one measurement per configuration
Run 2 — what we did not record
Run 2 has the engine commit, the quantization, the context length and both throughput figures, and nothing else: no driver version, no batch settings, no speculative-decoding settings.
Neither run has a saved command line. We did not keep one, so we are not going to reconstruct one after the fact — a plausible-looking command in a table is worse than a known gap. Treat both rows as measured once, not independently repeated: they carry the same "no second source" caveat as every other row on the board (see methodology).
If you want to reproduce either, the honest starting point is: llama-bench (or llama-server) on a 4090 D, Qwen3.8-27B at the quant above, the stated context, and for run 1 the MTP head with n-max=2. Your numbers will differ with the driver, the batch settings and the exact build.
Why this page exists
Because "we verified it" is a claim we cannot support for other people's measurements, and it would be dishonest to make an exception for our own. Publishing the two rows we did run — with the gaps visible — is the only version of that claim we can stand behind.
Sources
- Both rows: benchmark board (source column:
self) · RTX 4090 D runs - What the tiers mean: benchmark methodology
- The rest of the board's provenance: localmaxxing import and public X posts, per-row linked