Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
70

Wins 53 test cases

vs
Overall
26

Wins 6 test cases

DeepSeek: DeepSeek V4.1 Flash is clearly better. It wins 53 of the 59 test cases where they differ.

Learn to build evals like this in the AI Engineering Academy

Task by task

Average score on each task, best build on each side. A score is highlighted only when that model wins clearly more of the task's test cases. Hover a score for the test cases behind it.

TaskAB
Overall
70
26
Agents
63
49
Structured Extraction
87
13
Coding
27
42
Grounded Answers
87
33
Instruction Following
73
13
Business Logic
67
0
Safety
84
32

The race

2 lead changes

Score piled up over the 105 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

DeepSeek: DeepSeek V4.1 FlashSpace Bunny Alpha
0255075100255075100Perfect run7026

Every test case

105 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

53 won by DeepSeek: DeepSeek V4.1 Flash6 won by Space Bunny Alpha46 tied

Thinking spend

Only DeepSeek: DeepSeek V4.1 Flash shows visible thinking on this suite.

DeepSeek: DeepSeek V4.1 Flash
6.5Ktokens
Per test case, 93% of it thinking
Space Bunny Alpha
1.8Ktokens
Per test case, no visible thinking
TaskAB
Agents
3.7K75% thinking
1.3KNo thinking
Structured Extraction
5.2K93% thinking
1KNo thinking
Coding
10.7K98% thinking
3.9KNo thinking
Grounded Answers
4.9K92% thinking
860No thinking
Instruction Following
9.3K93% thinking
1.5KNo thinking
Business Logic
8.8K98% thinking
3.4KNo thinking
Safety
2.8K81% thinking
645No thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
108 tok/s
108 tok/s
Memory
—
—
Context
1M
1M
Cost per run
$0.2717
Not measured
Cost per passed case
$0.003881
Not measured
Quantization
FP8
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.