Head to head
Any two models from the board, on the same tasks and graders.
Ling-3.0 flash leads by 18 points and wins 9 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskALing-3.0 flashBLing 3.0 TinyLead
Overall
39
21
Tool Calling
67
37
Structured Output
37
20
RAG / Retrieval QA
30
4
Context Recall
89
56
Coding
11
0
Reasoning & Math
32
25
Instruction Following
7
4
Classification
45
17
Summarization
34
27
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, Ling-3.0 flash passes 20% of test cases and Ling 3.0 Tiny passes 8%.
TierALing-3.0 flashBLing 3.0 TinyLead
Baseline
75
56
Hard
46
19
Impossible
20
8
Thinking spend
Both models spend about the same on thinking per test case.
4.1Ktokens
Per test case, 55% of it thinking
4.4Ktokens
Per test case, 56% of it thinking
TaskALing-3.0 flashBLing 3.0 Tiny
Tool Calling
6.7K40% thinking
4.8K45% thinking
Structured Output
3K46% thinking
3.2K45% thinking
RAG / Retrieval QA
3.3K54% thinking
3.9K59% thinking
Context Recall
87459% thinking
1.8K66% thinking
Coding
8K79% thinking
8.2K79% thinking
Reasoning & Math
6.2K52% thinking
6.7K49% thinking
Instruction Following
3.2K54% thinking
3.8K51% thinking
Classification
2.8K50% thinking
3.4K46% thinking
Summarization
2K49% thinking
2.6K54% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecALing-3.0 flashBLing 3.0 Tiny
Speed
264 tok/s (better)
195 tok/s
Memory
—
—
Context
262.1K
262.1K
Cost per run
$0
$0
Quantization
—
—
Harness
openrouter
openrouter