Head to head

Any two models from the board, on the same tasks and graders.
Side A
Profile
Side B
Profile
Overall
60

Wins 9 tasks, 62 test cases

vs
Overall
39

Wins 0 tasks, 5 test cases

DeepSeek V4.1 Flash leads by 21 points and wins 9 of 9 tasks.

Task by task

DeepSeek V4.1 Flash wins 9 of 9

Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.

TaskAB
Overall
60
39
Tool Calling
78
67
Structured Output
52
37
RAG / Retrieval QA
63
30
Context Recall
100
89
Coding
41
11
Reasoning & Math
39
32
Instruction Following
46
7
Classification
55
45
Summarization
70
34

The race

No lead changes

Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.

DeepSeek V4.1 FlashLing-3.0 flash
025507510050100150200Perfect run5937

Every test case

238 test cases

One square per test case, in suite order. A square takes the color of the model that scored higher on it.

62 won by DeepSeek V4.1 Flash5 won by Ling-3.0 flash171 tied

Where it breaks

238 shared test cases

On the impossible tier, DeepSeek V4.1 Flash passes 44% of test cases and Ling-3.0 flash passes 20%.

TierAB
Baseline
36 test cases
94%
75%
Hard
69 test cases
67%
46%
Impossible
133 test cases
44%
20%

Tiers come from the suite. Baseline is fair, hard is adversarial, and impossible is built so that nothing solves it. Each number is the share of test cases the model fully passed.

Thinking spend

DeepSeek V4.1 Flash spends 1.4× as many thinking tokens per test case.

DeepSeek V4.1 Flash
3.2Ktokens
Per test case, 97% of it thinking
Ling-3.0 flash
4.1Ktokens
Per test case, 55% of it thinking
TaskAB
Tool Calling
3.7K84% thinking
6.7K40% thinking
Structured Output
2.3K99% thinking
3K46% thinking
RAG / Retrieval QA
2.8K100% thinking
3.3K54% thinking
Context Recall
59099% thinking
87459% thinking
Coding
6.3K96% thinking
8K79% thinking
Reasoning & Math
5.4K100% thinking
6.2K52% thinking
Instruction Following
2.9K99% thinking
3.2K54% thinking
Classification
2.5K100% thinking
2.8K50% thinking
Summarization
1.7K99% thinking
2K49% thinking

Average output tokens per test case, and how much of it was thinking.

What you'd run

The build behind each score. The better value in each row is highlighted.

SpecAB
Speed
215 tok/s
264 tok/s (better)
Memory
—
—
Context
1M (better)
262.1K
Cost per run
$0.7017
$0 (better)
Quantization
—
—
Harness
openrouter
openrouter

Faster, smaller, cheaper and more context count as better. Quantization and harness only describe the builds.