Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
70

Wins 74 test cases

vs
Overall
0

Wins 0 test cases

DeepSeek: DeepSeek V4.1 Flash is clearly better. It wins 74 of the 74 test cases where they differ.

Learn to build evals like this in the AI Engineering Academy

Task by task

Average score on each task, best build on each side. A score is highlighted only when that model wins clearly more of the task's test cases. Hover a score for the test cases behind it.

TaskAB
Overall
70
0
Agents
63
0
Structured Extraction
87
0
Coding
27
0
Grounded Answers
87
0
Instruction Following
73
0
Business Logic
67
0
Safety
84
0

The race

No lead changes

Score piled up over the 105 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

DeepSeek: DeepSeek V4.1 FlashLiquidAI: LFM2.5-2.6B (free)
0255075100255075100Perfect run700

Every test case

105 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

74 won by DeepSeek: DeepSeek V4.1 Flash0 won by LiquidAI: LFM2.5-2.6B (free)31 tied

Thinking spend

LiquidAI: LFM2.5-2.6B (free) spends 1.5× as many thinking tokens per test case.

DeepSeek: DeepSeek V4.1 Flash
6.5Ktokens
Per test case, 93% of it thinking
LiquidAI: LFM2.5-2.6B (free)
9.1Ktokens
Per test case, 96% of it thinking
TaskAB
Agents
3.7K75% thinking
16K98% thinking
Structured Extraction
5.2K93% thinking
6.6K97% thinking
Coding
10.7K98% thinking
8.1K93% thinking
Grounded Answers
4.9K92% thinking
6.2K98% thinking
Instruction Following
9.3K93% thinking
13.6K95% thinking
Business Logic
8.8K98% thinking
7.3K98% thinking
Safety
2.8K81% thinking
5.8K96% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
108 tok/s
229 tok/s (better)
Memory
—
—
Context
1M (better)
65.5K
Cost per run
$0.2717
Not measured
Cost per passed case
$0.003881
Not measured
Quantization
FP8
FP8
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.