Head to head
Any two models from the board, on the same tasks and graders.
DeepSeek V4.1 Flash leads by 7 points and wins 7 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskADeepSeek V4.1 FlashBDeepSeek V4 Flash 0731Lead
Overall
60
53
Tool Calling
78
85
Structured Output
52
37
RAG / Retrieval QA
63
59
Context Recall
100
94
Coding
41
52
Reasoning & Math
39
36
Instruction Following
46
18
Classification
55
45
Summarization
70
52
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, DeepSeek V4.1 Flash passes 44% of test cases and DeepSeek V4 Flash 0731 passes 38%.
TierADeepSeek V4.1 FlashBDeepSeek V4 Flash 0731Lead
Baseline
94
78
Hard
67
61
Impossible
44
38
Thinking spend
DeepSeek V4.1 Flash spends 1.2× as many thinking tokens per test case.
3.2Ktokens
Per test case, 97% of it thinking
3.5Ktokens
Per test case, 76% of it thinking
TaskADeepSeek V4.1 FlashBDeepSeek V4 Flash 0731
Tool Calling
3.7K84% thinking
4.1K52% thinking
Structured Output
2.3K99% thinking
2.8K78% thinking
RAG / Retrieval QA
2.8K100% thinking
2.9K85% thinking
Context Recall
59099% thinking
75787% thinking
Coding
6.3K96% thinking
6.1K86% thinking
Reasoning & Math
5.4K100% thinking
5.9K73% thinking
Instruction Following
2.9K99% thinking
3.4K71% thinking
Classification
2.5K100% thinking
2.7K81% thinking
Summarization
1.7K99% thinking
1.8K81% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecADeepSeek V4.1 FlashBDeepSeek V4 Flash 0731
Speed
215 tok/s (better)
65 tok/s
Memory
—
—
Context
1M
1M
Cost per run
$0.7017
$0.2742 (better)
Quantization
—
—
Harness
openrouter
openrouter