Instinct MI100 x4 vs RTX 3090
Best verified output tok/s per model (single-stream generation). Hover a value for the quant.
Instinct MI100 x41 verified run · best 132.4 tok/s
RTX 309016 verified runs · best 709.5 tok/s
| Model | Instinct MI100 x4 | RTX 3090 |
|---|
| LFM2.5-1.2B-Instruct | — | 709.5 Q4_K_M |
| Qwen3.5-0.8B-MTP | — | 476.1 Q6_K_XL |
| NVIDIA-Nemotron-3.5-Lightning-30B-A3B | — | 353.7 NVFP4 |
| LFM2.5-2.6B | — | 305.6 Q4_K_M |
| Qwen3.5-4B-MTP | — | 240.4 Q6_K_XL |
| Qwen3.8-27B | 132.4 INT8 | 164.6 EXL3-3.5bpw |
| Qwen3.5-9B-MTP | — | 164.3 Q6_K_XL |
| Qwen/Qwen3.8-27B | — | 136.8 AWQ-INT4 W4A16 |
| Qwen3.8-Flash-Next | — | 38.6 EXL3 2.50bpw |
| Ornstein3.6-27B-MTP-NSC-ACE-SABER | — | 32.9 Q4_K_M |
Head-to-head (models tested on both): Instinct MI100 x4 0 · RTX 3090 1 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.
Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board