Head to head
Any two models from the board, on the same tasks and graders.
DeepSeek V4.1 Flash leads by 12 points and wins 9 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskADeepSeek V4.1 FlashBQwen3.8 27BLead
Overall
60
48
Tool Calling
78
65
Structured Output
52
48
RAG / Retrieval QA
63
59
Context Recall
100
94
Coding
41
7
Reasoning & Math
39
32
Instruction Following
46
25
Classification
55
38
Summarization
70
62
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, DeepSeek V4.1 Flash passes 44% of test cases and Qwen3.8 27B passes 29%.
TierADeepSeek V4.1 FlashBQwen3.8 27BLead
Baseline
94
86
Hard
67
57
Impossible
44
29
Thinking spend
Both models spend about the same on thinking per test case.
3.2Ktokens
Per test case, 97% of it thinking
3.4Ktokens
Per test case, 98% of it thinking
TaskADeepSeek V4.1 FlashBQwen3.8 27B
Tool Calling
3.7K84% thinking
4.8K88% thinking
Structured Output
2.3K99% thinking
2.6K99% thinking
RAG / Retrieval QA
2.8K100% thinking
2.5K100% thinking
Context Recall
59099% thinking
50398% thinking
Coding
6.3K96% thinking
6.2K100% thinking
Reasoning & Math
5.4K100% thinking
5.4K100% thinking
Instruction Following
2.9K99% thinking
2.9K100% thinking
Classification
2.5K100% thinking
2.7K100% thinking
Summarization
1.7K99% thinking
1.9K99% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecADeepSeek V4.1 FlashBQwen3.8 27B
Speed
215 tok/s (better)
45 tok/s
Memory
—
—
Context
1M
—
Cost per run
$0.7017
$0 (better)
Quantization
—
—
Harness
openrouter
openai