Head to head
Any two models from the board, on the same tasks and graders.
Ling-3.0 flash leads by 7 points and wins 7 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskALaguna S 2.1BLing-3.0 flashLead
Overall
32
39
Tool Calling
60
67
Structured Output
30
37
RAG / Retrieval QA
19
30
Context Recall
50
89
Coding
54
11
Reasoning & Math
11
32
Instruction Following
4
7
Classification
17
45
Summarization
41
34
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, Laguna S 2.1 passes 19% of test cases and Ling-3.0 flash passes 20%.
TierALaguna S 2.1BLing-3.0 flashLead
Baseline
53
75
Hard
29
46
Impossible
19
20
Thinking spend
Ling-3.0 flash spends 1.2× as many thinking tokens per test case.
2.5Ktokens
Per test case, 76% of it thinking
4.1Ktokens
Per test case, 55% of it thinking
TaskALaguna S 2.1BLing-3.0 flash
Tool Calling
4.2K74% thinking
6.7K40% thinking
Structured Output
1.4K97% thinking
3K46% thinking
RAG / Retrieval QA
1.5K79% thinking
3.3K54% thinking
Context Recall
66683% thinking
87459% thinking
Coding
2.9K42% thinking
8K79% thinking
Reasoning & Math
5.6K87% thinking
6.2K52% thinking
Instruction Following
1.8K34% thinking
3.2K54% thinking
Classification
2.8K100% thinking
2.8K50% thinking
Summarization
51393% thinking
2K49% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecALaguna S 2.1BLing-3.0 flash
Speed
48 tok/s
264 tok/s (better)
Memory
—
—
Context
262.1K
262.1K
Cost per run
$0
$0
Quantization
—
—
Harness
openrouter
openrouter