Head to head
Any two models from the board, on the same tasks and graders.
DeepSeek: DeepSeek V4.1 Flash is clearly better. It wins 74 of the 74 test cases where they differ.
Task by task
Average score on each task, best build on each side. A score is highlighted only when that model wins clearly more of the task's test cases. Hover a score for the test cases behind it.
TaskADeepSeek: DeepSeek V4.1 FlashBLiquidAI: LFM2.5-2.6B (free)Lead
Overall
70
0
Agents
63
0
Structured Extraction
87
0
Coding
27
0
Grounded Answers
87
0
Instruction Following
73
0
Business Logic
67
0
Safety
84
0
The race
Score piled up over the 105 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Agents
Structured Extraction
Coding
Grounded Answers
Instruction Following
Business Logic
Safety
Thinking spend
LiquidAI: LFM2.5-2.6B (free) spends 1.5× as many thinking tokens per test case.
6.5Ktokens
Per test case, 93% of it thinking
9.1Ktokens
Per test case, 96% of it thinking
TaskADeepSeek: DeepSeek V4.1 FlashBLiquidAI: LFM2.5-2.6B (free)
Agents
3.7K75% thinking
16K98% thinking
Structured Extraction
5.2K93% thinking
6.6K97% thinking
Coding
10.7K98% thinking
8.1K93% thinking
Grounded Answers
4.9K92% thinking
6.2K98% thinking
Instruction Following
9.3K93% thinking
13.6K95% thinking
Business Logic
8.8K98% thinking
7.3K98% thinking
Safety
2.8K81% thinking
5.8K96% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecADeepSeek: DeepSeek V4.1 FlashBLiquidAI: LFM2.5-2.6B (free)
Speed
108 tok/s
229 tok/s (better)
Memory
—
—
Context
1M (better)
65.5K
Cost per run
$0.2717
Not measured
Cost per passed case
$0.003881
Not measured
Quantization
FP8
FP8
Harness
openrouter
openrouter