Head to head

Any two models from the board, on one scale.

Model A
Model B
Nanbeige4.2 3B

llama.cpp · bartowski · Q8_0 · 39.5 GB

overall

24

DeepSeek V4.1 Flash

openrouter · deepseekhosted

36vs60

DeepSeek V4.1 Flash leads by 24 overall · ahead on 2 of 2 tasks

A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.

the tape

Every case, one strip

55 cases

Nanbeige4.2 3B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
47TOOL+31 ▸78
not runSTRUCT·52
not runRAG·63
not runCTX·100
not runCODE·41
25REASON+14 ▸39
not runINSTR·46
not runCLASS·55
not runSUM·70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Nanbeige4.2 3B

5.4K tok/case · 66% thinking

DeepSeek V4.1 Flash

not run

5.5KTOOLnot run
5.3KREASONnot run

Only Nanbeige4.2 3B shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

33 tok/sspeed215 tok/s
39.5 GBmemory
193.8Kcontext1.0M
$1.5122 est.cost / run$0.7017
Q8_0quantization
llama.cppharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds