Head to head

Any two models from the board, on one scale.

Model A
Model B
Nemotron 3 Ultra

openrouter · nvidiahosted

overall

13

DeepSeek V4.1 Flash

openrouter · deepseekhosted

47vs60

DeepSeek V4.1 Flash leads by 13 overall · ahead on 8 of 9 tasks

the tape

Every case, one strip

7 cases

Nemotron 3 Ultra takes 0DeepSeek V4.1 Flash takes 00 even7 ran on one side only

one cell per case, in suite order · the brighter score wins the cell · hover for the case

tasks

Task by task

9 tasks
46TOOL+32 ▸78
41STRUCT+11 ▸52
22RAG+41 ▸63
78CTX+22 ▸100
74CODE◂ +3341
32REASON+7 ▸39
29INSTR+17 ▸46
41CLASS+14 ▸55
56SUM+14 ▸70

best build per side · same suite, same graders · hover a score for cases passed

tokens

Thinking spend

Nemotron 3 Ultra

1.7K tok/case · 53% thinking

DeepSeek V4.1 Flash

not run

1.6KCLASSnot run
1.8KSUMnot run

Only Nemotron 3 Ultra shows visible thinking on this suite

bar = average output tokens per case, shared scale · solid = answer · faded = thinking · hover a bar for the split

specs

What you'd run

21 tok/sspeed215 tok/s
memory
1Mcontext1.0M
$0cost / run$0.7017
quantization
openrouterharnessopenrouter

the lit value is the better spec — faster, smaller, cheaper, or more context · quantization & harness just describe the builds