DiggTop

Intel Arc Pro B70 vs RTX 3060 x2

Best verified output tok/s per model (single-stream generation). Hover a value for the quant.

Intel Arc Pro B7016 verified runs · best 283.1 tok/s
RTX 3060 x229 verified runs · best 198.1 tok/s
ModelIntel Arc Pro B70RTX 3060 x2
Laguna-XS-2.1283.1 Q5_K_M GGUF
Qwen3.6-35B-A3B248.3 GPTQ-4bit
Ornith-1.0-35B-B70-Turbo224.9 Q5_K_M GGUF
Nex-N2-mini-B70-Turbo224.3 Q5_K_M GGUF
SIQ-1-35B-B70-Turbo223.7 Q5_K_M GGUF
Qwen-AgentWorld-35B-A3B-B70-Turbo209.0 Q5_K_M GGUF
Qwen3-1.7B198.1 Q4_K_M
NVIDIA-Nemotron-3.5-Lightning-30B-A3B186.6 GPTQ-INT4-G64-sym-local+DFlash-B
LFM2.5-2.6B132.4 Q8_0
Ornith-1.5-35B-A3B131.5 Q4_K_M
MiniCPM5-1B125.9 Q4_K_M
Trinity-Mini109.0 Q4_K_M
Phi-4-mini-reasoning108.2 Q4_K_M
gemma-4-26B-A4B-it-qat95.3 Q4_K_M
Qwen3-30B-A3B-Thinking-250794.3 Q4_K_M
gpt-oss-20b93.4 Q4_K_M
OpenReasoning-Nemotron-7B68.0 Q4_K_M
WebWorld-8B62.3 Q4_K_M
gemma-4-26B-A4B-it62.1 UD-Q4_K_M
Gemma4-26B-A4B-QAT-Uncensored-HauhauCS-Balanced-MTP58.4 Q4_K_M
Carnice-9b56.1 Q4_K_M
Muse-Glimmer-30B29.2 Q4_K_XL56.0 UD-Q4_K_XL
Ornith-1.5-9B49.6 Q8_0
NVIDIA-Nemotron-Nano-12B-v240.4 Q4_K_M
gemma-3-12b-it39.4 Q4_K_M
Qwen3.6-27B-MTP36.0 Q8_0
Qwen3-14B35.8 Q4_K_M
Nemotron-Cascade-14B-Thinking35.7 Q4_K_M
Phi-4-reasoning35.4 Q4_K_M
WebWorld-14B35.4 Q4_K_M
InternVL3-14B35.1 Q4_K_M
DeepSeek-R1-Distill-Qwen-14B35.1 Q4_K_M
phi-434.2 Q4_K_M
ThinkingCap-Qwen3.6-27B27.9 Q6_K
Qwen3.8-27B27.8 Q4_K_M
DeepHermes-3-Mistral-24B-Preview24.5 Q4_K_M
reka-flash-323.9 Q4_K_M
mistralai_Mistral-Small-3.2-24B-Instruct-250621.8 Q4_K_M
Qwen3.6-27B18.9 IQ4_NL
Tess-4-27B18.8 Q4_K_M

Head-to-head (models tested on both): Intel Arc Pro B70 0 · RTX 3060 x2 1 · tied 0 · data as of Sep 23, 2026. Full rules: methodology.

Related reading: RTX 5090 vs 3090 · quantization · Value index · Full board