Head to head

Any two models from the board, on one scale.

Model A
Model B
Qwen3.5 4B

llama.cpp · unsloth · Q4_K_XL · 26.9 GB

overall

23

DeepSeek V4.1 Flash

openrouter · deepseekhosted

37vs60

DeepSeek V4.1 Flash leads by 23 overall · ahead on 2 of 2 tasks

A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.

the tape

Every case, one strip

55 cases

Qwen3.5 4B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
55TOOL+23 ▸78
not runSTRUCT·52
not runRAG·63
not runCTX·100
not runCODE·41
18REASON+21 ▸39
not runINSTR·46
not runCLASS·55
not runSUM·70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Qwen3.5 4B

10.6K tok/case · 80% thinking

DeepSeek V4.1 Flash

not run

14.9KTOOLnot run
6.5KREASONnot run

Only Qwen3.5 4B shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

85 tok/sspeed215 tok/s
26.9 GBmemory
262.1Kcontext1.0M
$1.7109 est.cost / run$0.7017
Q4_K_XLquantization
llama.cppharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds