Head to head
Any two models from the board, on the same tasks and graders.
MiMo-V2.6-Flash leads by 20 points and wins 7 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskALing 3.0 TinyBMiMo-V2.6-FlashLead
Overall
21
41
Tool Calling
37
66
Structured Output
20
47
RAG / Retrieval QA
4
33
Context Recall
56
100
Coding
0
31
Reasoning & Math
25
11
Instruction Following
4
14
Classification
17
17
Summarization
27
50
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, Ling 3.0 Tiny passes 8% of test cases and MiMo-V2.6-Flash passes 30%.
TierALing 3.0 TinyBMiMo-V2.6-FlashLead
Baseline
56
53
Hard
19
45
Impossible
8
30
Thinking spend
MiMo-V2.6-Flash spends 1.3× as many thinking tokens per test case.
4.4Ktokens
Per test case, 56% of it thinking
3.4Ktokens
Per test case, 95% of it thinking
TaskALing 3.0 TinyBMiMo-V2.6-Flash
Tool Calling
4.8K45% thinking
3.5K85% thinking
Structured Output
3.2K45% thinking
2.6K99% thinking
RAG / Retrieval QA
3.9K59% thinking
2.2K91% thinking
Context Recall
1.8K66% thinking
23195% thinking
Coding
8.2K79% thinking
7.2K95% thinking
Reasoning & Math
6.7K49% thinking
6.3K98% thinking
Instruction Following
3.8K51% thinking
2.9K99% thinking
Classification
3.4K46% thinking
2.9K98% thinking
Summarization
2.6K54% thinking
1.5K98% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecALing 3.0 TinyBMiMo-V2.6-Flash
Speed
195 tok/s (better)
69 tok/s
Memory
—
—
Context
262.1K
1M (better)
Cost per run
$0 (better)
$0.4321
Quantization
—
—
Harness
openrouter
openrouter