Head to head
Any two models from the board, on one scale.
Model A
Model B
32vs60
DeepSeek V4.1 Flash leads by 28 overall · ahead on 8 of 9 tasks
the tape
Every case, one strip
156 casesLaguna S 2.1 takes 0DeepSeek V4.1 Flash takes 00 even156 ran on one side only
one cell per case, in suite order · the brighter score wins the cell · hover for the case
tasks
Task by task
9 tasks60TOOL+18 ▸78
30STRUCT+22 ▸52
19RAG+44 ▸63
50CTX+50 ▸100
54CODE◂ +1341
11REASON+28 ▸39
4INSTR+42 ▸46
17CLASS+38 ▸55
41SUM+29 ▸70
best build per side · same suite, same graders · hover a score for cases passed
tokens
Thinking spend
Laguna S 2.1
1.8K tok/case · 70% thinking
DeepSeek V4.1 Flash
not run
3.4KTOOLnot run—
1.1KSTRUCTnot run—
1.8KRAGnot run—
776CTXnot run—
2.2KCODEnot run—
3.7KREASONnot run—
1.1KINSTRnot run—
1.6KCLASSnot run—
70SUMnot run—
Only Laguna S 2.1 shows visible thinking on this suite
bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split
specs
What you'd run
48 tok/sspeed215 tok/s
—memory—
262.1Kcontext1.0M
$0cost / run$0.7017
—quantization—
openrouterharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds