Head to head
Any two models from the board, on one scale.
Model A
Model B
41vs60
DeepSeek V4.1 Flash leads by 19 overall · ahead on 8 of 9 tasks
tasks
Task by task
9 tasks66TOOL+12 ▸78
47STRUCT+5 ▸52
33RAG+30 ▸63
100CTXeven100
31CODE+10 ▸41
11REASON+28 ▸39
14INSTR+32 ▸46
17CLASS+38 ▸55
50SUM+20 ▸70
best build per side · same suite, same graders · hover a score for cases passed
specs
What you'd run
69 tok/sspeed215 tok/s
—memory—
1.0Mcontext1.0M
$0.4321cost / run$0.7017
—quantization—
openrouterharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds