Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
39

Wins 3 tasks, 23 test cases

vs
Overall
41

Wins 6 tasks, 28 test cases

MiMo-V2.6-Flash leads by 2 points and wins 6 of 9 tasks.

Task by task

MiMo-V2.6-Flash wins 6 of 9

Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.

TaskAB
Overall
39
41
Tool Calling
67
66
Structured Output
37
47
RAG / Retrieval QA
30
33
Context Recall
89
100
Coding
11
31
Reasoning & Math
32
11
Instruction Following
7
14
Classification
45
17
Summarization
34
50

The race

6 lead changes

Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

Ling-3.0 flashMiMo-V2.6-Flash
025507510050100150200Perfect run3738

Every test case

238 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

23 won by Ling-3.0 flash28 won by MiMo-V2.6-Flash187 tied

Where it breaks

238 shared test cases

On the impossible tier, Ling-3.0 flash passes 20% of test cases and MiMo-V2.6-Flash passes 30%.

TierAB
Baseline
36 test cases
75%
53%
Hard
69 test cases
46%
45%
Impossible
133 test cases
20%
30%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

MiMo-V2.6-Flash spends 1.4× as many thinking tokens per test case.

Ling-3.0 flash
4.1Ktokens
Per test case, 55% of it thinking
MiMo-V2.6-Flash
3.4Ktokens
Per test case, 95% of it thinking
TaskAB
Tool Calling
6.7K40% thinking
3.5K85% thinking
Structured Output
3K46% thinking
2.6K99% thinking
RAG / Retrieval QA
3.3K54% thinking
2.2K91% thinking
Context Recall
87459% thinking
23195% thinking
Coding
8K79% thinking
7.2K95% thinking
Reasoning & Math
6.2K52% thinking
6.3K98% thinking
Instruction Following
3.2K54% thinking
2.9K99% thinking
Classification
2.8K50% thinking
2.9K98% thinking
Summarization
2K49% thinking
1.5K98% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
264 tok/s (better)
69 tok/s
Memory
—
—
Context
262.1K
1M (better)
Cost per run
$0 (better)
$0.4321
Quantization
—
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.