Head to head

Any two models from the board, on one scale.

Model A
Model B
DeepSeek V4 Flash 0731

openrouter · deepseekhosted

overall

7

DeepSeek V4.1 Flash

openrouter · deepseekhosted

53vs60

DeepSeek V4.1 Flash leads by 7 overall · ahead on 7 of 9 tasks

the tape

Every case, one strip

238 cases

DeepSeek V4 Flash 0731 takes 0DeepSeek V4.1 Flash takes 00 even238 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
85TOOL◂ +778
37STRUCT+15 ▸52
59RAG+4 ▸63
94CTX+6 ▸100
52CODE◂ +1141
36REASON+3 ▸39
18INSTR+28 ▸46
45CLASS+10 ▸55
52SUM+18 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

DeepSeek V4 Flash 0731

3.5K tok/case · 76% thinking

DeepSeek V4.1 Flash

not run

4.1KTOOLnot run
2.8KSTRUCTnot run
2.9KRAGnot run
757CTXnot run
6.1KCODEnot run
5.9KREASONnot run
3.4KINSTRnot run
2.7KCLASSnot run
1.8KSUMnot run

Only DeepSeek V4 Flash 0731 shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

65 tok/sspeed215 tok/s
memory
1.0Mcontext1.0M
$0.2742cost / run$0.7017
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds