Head to head

Any two models from the board, on one scale.

Model A
Model B
Gemma 4 12B

llama.cpp · unsloth · Q4_K_XL · 26.1 GB

overall

32

DeepSeek V4.1 Flash

openrouter · deepseekhosted

28vs60

DeepSeek V4.1 Flash leads by 32 overall · ahead on 2 of 2 tasks

A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.

the tape

Every case, one strip

55 cases

Gemma 4 12B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
30TOOL+48 ▸78
not runSTRUCT·52
not runRAG·63
not runCTX·100
not runCODE·41
25REASON+14 ▸39
not runINSTR·46
not runCLASS·55
not runSUM·70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Gemma 4 12B

4.7K tok/case · 78% thinking

DeepSeek V4.1 Flash

not run

4.3KTOOLnot run
5.1KREASONnot run

Only Gemma 4 12B shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

58 tok/sspeed215 tok/s
26.1 GBmemory
262.1Kcontext1.0M
$1.0308 est.cost / run$0.7017
Q4_K_XLquantization
llama.cppharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds