Head to head
Space Bunny Alpha is clearly better. It wins 56 of the 59 test cases where they differ.
Task by task
Average score on each task, best build on each side. A score is highlighted only when that model wins clearly more of the task's test cases. Hover a score for the test cases behind it.
The race
Score piled up over the 105 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Where it breaks
On the impossible tier, LiquidAI: LFM2.5-2.6B (free) passes 2% of test cases and Space Bunny Alpha passes 52%.
Thinking spend
Only LiquidAI: LFM2.5-2.6B (free) shows visible thinking on this suite.
What you'd run
The build behind each score. The better value in each row is highlighted.