Head to head

Any two models from the board, on one scale.

Model A
Model B
Laguna S 2.1

openrouter · poolsidehosted

overall

28

DeepSeek V4.1 Flash

openrouter · deepseekhosted

32vs60

DeepSeek V4.1 Flash leads by 28 overall · ahead on 8 of 9 tasks

the tape

Every case, one strip

156 cases

Laguna S 2.1 takes 0DeepSeek V4.1 Flash takes 00 even156 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
60TOOL+18 ▸78
30STRUCT+22 ▸52
19RAG+44 ▸63
50CTX+50 ▸100
54CODE◂ +1341
11REASON+28 ▸39
4INSTR+42 ▸46
17CLASS+38 ▸55
41SUM+29 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Laguna S 2.1

1.8K tok/case · 70% thinking

DeepSeek V4.1 Flash

not run

3.4KTOOLnot run
1.1KSTRUCTnot run
1.8KRAGnot run
776CTXnot run
2.2KCODEnot run
3.7KREASONnot run
1.1KINSTRnot run
1.6KCLASSnot run
70SUMnot run

Only Laguna S 2.1 shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

48 tok/sspeed215 tok/s
memory
262.1Kcontext1.0M
$0cost / run$0.7017
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds