Head to head
Any two models from the board, on one scale.
Model A
Model B
47vs60
DeepSeek V4.1 Flash leads by 13 overall · ahead on 8 of 9 tasks
the tape
Every case, one strip
7 casesNemotron 3 Ultra takes 0DeepSeek V4.1 Flash takes 00 even7 ran on one side only
one cell per case, in suite order · the brighter score wins the cell · hover for the case
tasks
Task by task
9 tasks46TOOL+32 ▸78
41STRUCT+11 ▸52
22RAG+41 ▸63
78CTX+22 ▸100
74CODE◂ +3341
32REASON+7 ▸39
29INSTR+17 ▸46
41CLASS+14 ▸55
56SUM+14 ▸70
best build per side · same suite, same graders · hover a score for cases passed
tokens
Thinking spend
Nemotron 3 Ultra
1.7K tok/case · 53% thinking
DeepSeek V4.1 Flash
not run
1.6KCLASSnot run—
1.8KSUMnot run—
Only Nemotron 3 Ultra shows visible thinking on this suite
bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split
specs
What you'd run
21 tok/sspeed215 tok/s
—memory—
1Mcontext1.0M
$0cost / run$0.7017
—quantization—
openrouterharnessopenrouter
the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds