Head to head
Any two models from the board, on the same tasks and graders.
Qwen3.8 Flash leads by 7 points and wins 6 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskANemotron 3 UltraBQwen3.8 FlashLead
Overall
47
54
Tool Calling
46
79
Structured Output
41
47
RAG / Retrieval QA
22
41
Context Recall
78
89
Coding
74
75
Reasoning & Math
32
32
Instruction Following
29
25
Classification
41
38
Summarization
56
60
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, Nemotron 3 Ultra passes 26% of test cases and Qwen3.8 Flash passes 35%.
TierANemotron 3 UltraBQwen3.8 FlashLead
Baseline
78
89
Hard
55
57
Impossible
26
35
Thinking spend
Qwen3.8 Flash spends 1.5× as many thinking tokens per test case.
2.6Ktokens
Per test case, 56% of it thinking
3Ktokens
Per test case, 70% of it thinking
TaskANemotron 3 UltraBQwen3.8 Flash
Tool Calling
4.3K47% thinking
6.7K66% thinking
Structured Output
1.5K47% thinking
1.9K67% thinking
RAG / Retrieval QA
1.9K62% thinking
2.4K66% thinking
Context Recall
1K65% thinking
55595% thinking
Coding
5K69% thinking
5.2K71% thinking
Reasoning & Math
3.6K54% thinking
4.3K73% thinking
Instruction Following
2.2K47% thinking
2.3K69% thinking
Classification
1.6K57% thinking
1.8K72% thinking
Summarization
1.4K59% thinking
1.4K70% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecANemotron 3 UltraBQwen3.8 Flash
Speed
21 tok/s
64 tok/s (better)
Memory
—
—
Context
1M
1M
Cost per run
$0 (better)
$0.5705
Quantization
—
—
Harness
openrouter
openrouter