Head to head
Any two models from the board, on the same tasks and graders.
Dots3-Note Preview leads by 8 points and wins 7 of 9 tasks.
Task by task
Average score on each task, best build on each side. The higher score is highlighted. Hover a score for the test cases behind it.
TaskADots3-Note PreviewBLaguna S 2.1Lead
Overall
40
32
Tool Calling
66
60
Structured Output
30
30
RAG / Retrieval QA
22
19
Context Recall
89
50
Coding
4
54
Reasoning & Math
32
11
Instruction Following
25
4
Classification
31
17
Summarization
58
41
The race
Score piled up over the 238 test cases both models ran, as a share of the suite's maximum. The dashed line is a perfect run.
Every test case
One square per test case, in suite order. A square takes the color of the model that scored higher on it.
Tool Calling
Structured Output
RAG / Retrieval QA
Context Recall
Coding
Reasoning & Math
Instruction Following
Classification
Summarization
Where it breaks
On the impossible tier, Dots3-Note Preview passes 23% of test cases and Laguna S 2.1 passes 19%.
TierADots3-Note PreviewBLaguna S 2.1Lead
Baseline
86
53
Hard
38
29
Impossible
23
19
Thinking spend
Dots3-Note Preview spends 1.5× as many thinking tokens per test case.
4.5Ktokens
Per test case, 62% of it thinking
2.5Ktokens
Per test case, 76% of it thinking
TaskADots3-Note PreviewBLaguna S 2.1
Tool Calling
8.8K52% thinking
4.2K74% thinking
Structured Output
2.9K49% thinking
1.4K97% thinking
RAG / Retrieval QA
3.4K58% thinking
1.5K79% thinking
Context Recall
1.7K65% thinking
66683% thinking
Coding
8K88% thinking
2.9K42% thinking
Reasoning & Math
6.2K57% thinking
5.6K87% thinking
Instruction Following
3.4K59% thinking
1.8K34% thinking
Classification
3.1K58% thinking
2.8K100% thinking
Summarization
2.5K61% thinking
51393% thinking
What you'd run
The build behind each score. The better value in each row is highlighted.
SpecADots3-Note PreviewBLaguna S 2.1
Speed
75 tok/s (better)
48 tok/s
Memory
—
—
Context
512K (better)
262.1K
Cost per run
$0
$0
Quantization
—
—
Harness
openrouter
openrouter