Head to head

Any two models from the board, on one scale.

Model A
Model B
Ling-3.0 flash

openrouter · inclusionaihosted

overall

21

DeepSeek V4.1 Flash

openrouter · deepseekhosted

39vs60

DeepSeek V4.1 Flash leads by 21 overall · ahead on 9 of 9 tasks

the tape

Every case, one strip

238 cases

Ling-3.0 flash takes 0DeepSeek V4.1 Flash takes 00 even238 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
67TOOL+11 ▸78
37STRUCT+15 ▸52
30RAG+33 ▸63
89CTX+11 ▸100
11CODE+30 ▸41
32REASON+7 ▸39
7INSTR+39 ▸46
45CLASS+10 ▸55
34SUM+36 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Ling-3.0 flash

4.1K tok/case · 55% thinking

DeepSeek V4.1 Flash

not run

6.7KTOOLnot run
3.0KSTRUCTnot run
3.3KRAGnot run
874CTXnot run
8.0KCODEnot run
6.2KREASONnot run
3.2KINSTRnot run
2.8KCLASSnot run
2.0KSUMnot run

Only Ling-3.0 flash shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

264 tok/sspeed215 tok/s
memory
262.1Kcontext1.0M
$0cost / run$0.7017
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds