Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
26

Wins 3 test cases

vs
Overall
73

Wins 56 test cases

Space Bunny Alpha is clearly better. It wins 56 of the 59 test cases where they differ.

Task by task

Average score on each task, best build on each side. A score is highlighted only when that model wins clearly more of the task's test cases. Hover a score for the test cases behind it.

TaskAB
Overall
26
73
Agents
16
63
Structured Extraction
9
100
Coding
22
87
Grounded Answers
38
80
Instruction Following
27
67
Business Logic
7
40
Safety
60
73

The race

No lead changes

Score piled up over the 105 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

LiquidAI: LFM2.5-2.6B (free)Space Bunny Alpha
0255075100255075100Perfect run2673

Every test case

105 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

3 won by LiquidAI: LFM2.5-2.6B (free)56 won by Space Bunny Alpha46 tied

Where it breaks

105 shared test cases

On the impossible tier, LiquidAI: LFM2.5-2.6B (free) passes 2% of test cases and Space Bunny Alpha passes 52%.

TierAB
Baseline
28 test cases
57%
89%
Hard
35 test cases
26%
77%
Impossible
42 test cases
2%
52%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

Only LiquidAI: LFM2.5-2.6B (free) shows visible thinking on this suite.

LiquidAI: LFM2.5-2.6B (free)
2.7Ktokens
Per test case, 92% of it thinking
Space Bunny Alpha
399tokens
Per test case, no visible thinking
TaskAB
Agents
3.6K89% thinking
520No thinking
Structured Extraction
2.1K90% thinking
259No thinking
Coding
4.4K96% thinking
773No thinking
Grounded Answers
1K86% thinking
74No thinking
Instruction Following
2K94% thinking
381No thinking
Business Logic
4.7K99% thinking
422No thinking
Safety
86762% thinking
363No thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
149 tok/s (better)
108 tok/s
Memory
—
—
Context
65.5K
1M (better)
Cost per run
Not measured
Not measured
Cost per passed case
Not measured
Not measured
Quantization
FP8
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.