Head to head

Any two models from the board, on one scale.

Model A
Model B
Ling 3.0 Tiny

openrouter · inclusionaihosted

overall

39

DeepSeek V4.1 Flash

openrouter · deepseekhosted

21vs60

DeepSeek V4.1 Flash leads by 39 overall · ahead on 9 of 9 tasks

the tape

Every case, one strip

84 cases

Ling 3.0 Tiny takes 0DeepSeek V4.1 Flash takes 00 even84 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
37TOOL+41 ▸78
20STRUCT+32 ▸52
4RAG+59 ▸63
56CTX+44 ▸100
0CODE+41 ▸41
25REASON+14 ▸39
4INSTR+42 ▸46
17CLASS+38 ▸55
27SUM+43 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Ling 3.0 Tiny

3.9K tok/case · 62% thinking

DeepSeek V4.1 Flash

not run

4.6KTOOLnot run
2.7KSTRUCTnot run
3.5KRAGnot run
1.9KCTXnot run
8.2KCODEnot run
5.3KREASONnot run
3.5KINSTRnot run
2.9KCLASSnot run
1.5KSUMnot run

Only Ling 3.0 Tiny shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

195 tok/sspeed215 tok/s
memory
262.1Kcontext1.0M
$0cost / run$0.7017
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds