Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
47

Wins 3 tasks, 35 test cases

vs
Overall
48

Wins 5 tasks, 33 test cases

Qwen3.8 27B leads by 1 point and wins 5 of 9 tasks.

Task by task

Qwen3.8 27B wins 5 of 9

Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.

TaskAB
Overall
47
48
Tool Calling
46
65
Structured Output
41
48
RAG / Retrieval QA
22
59
Context Recall
78
94
Coding
74
7
Reasoning & Math
32
32
Instruction Following
29
25
Classification
41
38
Summarization
56
62

The race

3 lead changes

Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

Nemotron 3 UltraQwen3.8 27B
025507510050100150200Perfect run4546

Every test case

238 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

35 won by Nemotron 3 Ultra33 won by Qwen3.8 27B170 tied

Where it breaks

238 shared test cases

On the impossible tier, Nemotron 3 Ultra passes 26% of test cases and Qwen3.8 27B passes 29%.

TierAB
Baseline
36 test cases
78%
86%
Hard
69 test cases
55%
57%
Impossible
133 test cases
26%
29%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

Qwen3.8 27B spends 2.3× as many thinking tokens per test case.

Nemotron 3 Ultra
2.6Ktokens
Per test case, 56% of it thinking
Qwen3.8 27B
3.4Ktokens
Per test case, 98% of it thinking
TaskAB
Tool Calling
4.3K47% thinking
4.8K88% thinking
Structured Output
1.5K47% thinking
2.6K99% thinking
RAG / Retrieval QA
1.9K62% thinking
2.5K100% thinking
Context Recall
1K65% thinking
50398% thinking
Coding
5K69% thinking
6.2K100% thinking
Reasoning & Math
3.6K54% thinking
5.4K100% thinking
Instruction Following
2.2K47% thinking
2.9K100% thinking
Classification
1.6K57% thinking
2.7K100% thinking
Summarization
1.4K59% thinking
1.9K99% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
21 tok/s
45 tok/s (better)
Memory
—
—
Context
1M
—
Cost per run
$0
$0
Quantization
—
—
Harness
openrouter
openai

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.