Head to head

Any two models from the board, on one scale.

Model A
Model B
Qwen3.5 9B

llama.cpp · unsloth · Q8_K_XL · 32.7 GB

overall

25

DeepSeek V4.1 Flash

openrouter · deepseekhosted

35vs60

DeepSeek V4.1 Flash leads by 25 overall · ahead on 2 of 2 tasks

A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.

the tape

Every case, one strip

55 cases

Qwen3.5 9B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
49TOOL+29 ▸78
not runSTRUCT·52
not runRAG·63
not runCTX·100
not runCODE·41
21REASON+18 ▸39
not runINSTR·46
not runCLASS·55
not runSUM·70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Qwen3.5 9B

12.4K tok/case · 88% thinking

DeepSeek V4.1 Flash

not run

20.0KTOOLnot run
5.2KREASONnot run

Only Qwen3.5 9B shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

24 tok/sspeed215 tok/s
32.7 GBmemory
262.1Kcontext1.0M
$7.2603 est.cost / run$0.7017
Q8_K_XLquantization
llama.cppharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds