Head to head

Any two models from the board, on one scale.

Model A
Model B
LFM2.5 2.6B

llama.cpp · liquidai · Q8_0 · 13.8 GB

overall

44

DeepSeek V4.1 Flash

openrouter · deepseekhosted

16vs60

DeepSeek V4.1 Flash leads by 44 overall · ahead on 2 of 2 tasks

A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.

the tape

Every case, one strip

55 cases

LFM2.5 2.6B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
18TOOL+60 ▸78
not runSTRUCT·52
not runRAG·63
not runCTX·100
not runCODE·41
14REASON+25 ▸39
not runINSTR·46
not runCLASS·55
not runSUM·70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

LFM2.5 2.6B

9.9K tok/case · 94% thinking

DeepSeek V4.1 Flash

not run

17.4KTOOLnot run
2.7KREASONnot run

Only LFM2.5 2.6B shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

84 tok/sspeed215 tok/s
13.8 GBmemory
128Kcontext1.0M
$1.1119 est.cost / run$0.7017
Q8_0quantization
llama.cppharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds