Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
21

Wins 0 tasks, 1 test case

vs
Overall
54

Wins 9 tasks, 91 test cases

Qwen3.8 Flash leads by 33 points and wins 9 of 9 tasks.

Task by task

Qwen3.8 Flash wins 9 of 9

Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.

TaskAB
Overall
21
54
Tool Calling
37
79
Structured Output
20
47
RAG / Retrieval QA
4
41
Context Recall
56
89
Coding
0
75
Reasoning & Math
25
32
Instruction Following
4
25
Classification
17
38
Summarization
27
60

The race

No lead changes

Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

Ling 3.0 TinyQwen3.8 Flash
025507510050100150200Perfect run2052

Every test case

238 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

1 won by Ling 3.0 Tiny91 won by Qwen3.8 Flash146 tied

Where it breaks

238 shared test cases

On the impossible tier, Ling 3.0 Tiny passes 8% of test cases and Qwen3.8 Flash passes 35%.

TierAB
Baseline
36 test cases
56%
89%
Hard
69 test cases
19%
57%
Impossible
133 test cases
8%
35%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

Ling 3.0 Tiny spends 1.2× as many thinking tokens per test case.

Ling 3.0 Tiny
4.4Ktokens
Per test case, 56% of it thinking
Qwen3.8 Flash
3Ktokens
Per test case, 70% of it thinking
TaskAB
Tool Calling
4.8K45% thinking
6.7K66% thinking
Structured Output
3.2K45% thinking
1.9K67% thinking
RAG / Retrieval QA
3.9K59% thinking
2.4K66% thinking
Context Recall
1.8K66% thinking
55595% thinking
Coding
8.2K79% thinking
5.2K71% thinking
Reasoning & Math
6.7K49% thinking
4.3K73% thinking
Instruction Following
3.8K51% thinking
2.3K69% thinking
Classification
3.4K46% thinking
1.8K72% thinking
Summarization
2.6K54% thinking
1.4K70% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
195 tok/s (better)
64 tok/s
Memory
—
—
Context
262.1K
1M (better)
Cost per run
$0 (better)
$0.5705
Quantization
—
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.