Head to head

Any two models from the board, on one scale.

Model A
Model B
DeepSeek V4.1 Flash

openrouter · deepseekhosted

overall

6

Qwen3.8 Flash

openrouter · qwenhosted

60vs54

DeepSeek V4.1 Flash leads by 6 overall · ahead on 7 of 9 tasks

tasks

Task by task

9 tasks
78TOOL+1 ▸79
52STRUCT◂ +547
63RAG◂ +2241
100CTX◂ +1189
41CODE+34 ▸75
39REASON◂ +732
46INSTR◂ +2125
55CLASS◂ +1738
70SUM◂ +1060

best build per side · same suite, same graders · hover a score for cases passed

specs

What you'd run

215 tok/sspeed64 tok/s
memory
1.0Mcontext1M
$0.7017cost / run$0.5705
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds