Head to head
Any two models from the board, on one scale.
Model A
Model B
60vs54
DeepSeek V4.1 Flash leads by 6 overall · ahead on 7 of 9 tasks
tasks
Task by task
9 tasks78TOOL+1 ▸79
52STRUCT◂ +547
63RAG◂ +2241
100CTX◂ +1189
41CODE+34 ▸75
39REASON◂ +732
46INSTR◂ +2125
55CLASS◂ +1738
70SUM◂ +1060
best build per side · same suite, same graders · hover a score for cases passed
specs
What you'd run
215 tok/sspeed64 tok/s
—memory—
1.0Mcontext1M
$0.7017cost / run$0.5705
—quantization—
openrouterharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds