Head to head
Any two models from the board, on the same tasks and graders.
Nemotron 3 Ultra leads by 6 points and wins 5 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskAMiMo-V2.6-FlashBNemotron 3 UltraLead
Overall
41
47
Tool Calling
66
46
Structured Output
47
41
RAG / Retrieval QA
33
22
Context Recall
100
78
Coding
31
74
Reasoning & Math
11
32
Instruction Following
14
29
Classification
17
41
Summarization
50
56
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, MiMo-V2.6-Flash passes 30% of test cases and Nemotron 3 Ultra passes 26%.
TierAMiMo-V2.6-FlashBNemotron 3 UltraLead
Baseline
53
78
Hard
45
55
Impossible
30
26
Thinking spend
MiMo-V2.6-Flash spends 2.2× as many thinking tokens per test case.
3.4Ktokens
Per test case, 95% of it thinking
2.6Ktokens
Per test case, 56% of it thinking
TaskAMiMo-V2.6-FlashBNemotron 3 Ultra
Tool Calling
3.5K85% thinking
4.3K47% thinking
Structured Output
2.6K99% thinking
1.5K47% thinking
RAG / Retrieval QA
2.2K91% thinking
1.9K62% thinking
Context Recall
23195% thinking
1K65% thinking
Coding
7.2K95% thinking
5K69% thinking
Reasoning & Math
6.3K98% thinking
3.6K54% thinking
Instruction Following
2.9K99% thinking
2.2K47% thinking
Classification
2.9K98% thinking
1.6K57% thinking
Summarization
1.5K98% thinking
1.4K59% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecAMiMo-V2.6-FlashBNemotron 3 Ultra
Speed
69 tok/s (better)
21 tok/s
Memory
—
—
Context
1M
1M
Cost per run
$0.4321
$0 (better)
Quantization
—
—
Harness
openrouter
openrouter