Head to head
Any two models from the board, on one scale.
llama.cpp · liquidai · Q8_0 · 13.8 GB
44 ▸
openrouter · deepseekhosted
DeepSeek V4.1 Flash leads by 44 overall · ahead on 2 of 2 tasks
A measured average only covers the tasks run for that model; use the task-by-task rows below for a like-for-like comparison.
Every case, one strip
55 casesLFM2.5 2.6B takes 0DeepSeek V4.1 Flash takes 00 even55 ran on one side only
one cell per case, in suite order · the brighter score wins the cell · hover for the case
Task by task
9 tasksbest build per side · same suite, same graders · hover a score for cases passed
Thinking spend
LFM2.5 2.6B
9.9K tok/case · 94% thinking
DeepSeek V4.1 Flash
not run
Only LFM2.5 2.6B shows visible thinking on this suite
bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split
What you'd run
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds