Head to head
Any two models from the board, on one scale.
Model A
Model B
48vs60
DeepSeek V4.1 Flash leads by 12 overall · ahead on 9 of 9 tasks
the tape
Every case, one strip
2 casesQwen3.8 27B takes 0DeepSeek V4.1 Flash takes 00 even2 ran on one side only
one cell per case, in suite order · the brighter score wins the cell · hover for the case
tasks
Task by task
9 tasks65TOOL+13 ▸78
48STRUCT+4 ▸52
59RAG+4 ▸63
94CTX+6 ▸100
7CODE+34 ▸41
32REASON+7 ▸39
25INSTR+21 ▸46
38CLASS+17 ▸55
62SUM+8 ▸70
best build per side · same suite, same graders · hover a score for cases passed
tokens
Thinking spend
Qwen3.8 27B
0 tok/case · no visible thinking
DeepSeek V4.1 Flash
not run
0REASONnot run—
Neither model shows visible thinking on this suite
bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split
specs
What you'd run
45 tok/sspeed215 tok/s
—memory—
—context1.0M
$0cost / run$0.7017
—quantization—
openaiharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds