Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
40

Wins 9 tasks, 54 test cases

vs
Overall
21

Wins 0 tasks, 4 test cases

Dots3-Note Preview leads by 19 points and wins 9 of 9 tasks.

Task by task

Dots3-Note Preview wins 9 of 9

Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.

TaskAB
Overall
40
21
Tool Calling
66
37
Structured Output
30
20
RAG / Retrieval QA
22
4
Context Recall
89
56
Coding
4
0
Reasoning & Math
32
25
Instruction Following
25
4
Classification
31
17
Summarization
58
27

The race

No lead changes

Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

Dots3-Note PreviewLing 3.0 Tiny
025507510050100150200Perfect run3820

Every test case

238 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

54 won by Dots3-Note Preview4 won by Ling 3.0 Tiny180 tied

Where it breaks

238 shared test cases

On the impossible tier, Dots3-Note Preview passes 23% of test cases and Ling 3.0 Tiny passes 8%.

TierAB
Baseline
36 test cases
86%
56%
Hard
69 test cases
38%
19%
Impossible
133 test cases
23%
8%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

Both models spend about the same on thinking per test case.

Dots3-Note Preview
4.5Ktokens
Per test case, 62% of it thinking
Ling 3.0 Tiny
4.4Ktokens
Per test case, 56% of it thinking
TaskAB
Tool Calling
8.8K52% thinking
4.8K45% thinking
Structured Output
2.9K49% thinking
3.2K45% thinking
RAG / Retrieval QA
3.4K58% thinking
3.9K59% thinking
Context Recall
1.7K65% thinking
1.8K66% thinking
Coding
8K88% thinking
8.2K79% thinking
Reasoning & Math
6.2K57% thinking
6.7K49% thinking
Instruction Following
3.4K59% thinking
3.8K51% thinking
Classification
3.1K58% thinking
3.4K46% thinking
Summarization
2.5K61% thinking
2.6K54% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
75 tok/s
195 tok/s (better)
Memory
—
—
Context
512K (better)
262.1K
Cost per run
$0
$0
Quantization
—
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.