Head to head

Any two models from the board, on one scale.

Model A
Model B
Qwen3.8 27B

openaihosted

overall

12

DeepSeek V4.1 Flash

openrouter · deepseekhosted

48vs60

DeepSeek V4.1 Flash leads by 12 overall · ahead on 9 of 9 tasks

the tape

Every case, one strip

2 cases

Qwen3.8 27B takes 0DeepSeek V4.1 Flash takes 00 even2 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
65TOOL+13 ▸78
48STRUCT+4 ▸52
59RAG+4 ▸63
94CTX+6 ▸100
7CODE+34 ▸41
32REASON+7 ▸39
25INSTR+21 ▸46
38CLASS+17 ▸55
62SUM+8 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Qwen3.8 27B

0 tok/case · no visible thinking

DeepSeek V4.1 Flash

not run

0REASONnot run

Neither model shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

45 tok/sspeed215 tok/s
memory
context1.0M
$0cost / run$0.7017
quantization
openaiharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds