Head to head
Any two models from the board, on one scale.
Model A
Model B
54vs60
DeepSeek V4.1 Flash leads by 6 overall · ahead on 7 of 9 tasks
tasks
Task by task
9 tasks79TOOL◂ +178
47STRUCT+5 ▸52
41RAG+22 ▸63
89CTX+11 ▸100
75CODE◂ +3441
32REASON+7 ▸39
25INSTR+21 ▸46
38CLASS+17 ▸55
60SUM+10 ▸70
best build per side · same suite, same graders · hover a score for cases passed
specs
What you'd run
64 tok/sspeed215 tok/s
—memory—
1Mcontext1.0M
$0.5705cost / run$0.7017
—quantization—
openrouterharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds