Which open-weight model should you run?

Real build-tasks, ranked for the hardware you have. The hosted frontier is the reference line — every score reproducible.

the headline matchupQwen3.5 4B 37 vs DeepSeek V4.1 Flash 60

the full matchup →

14 models·238 cases·9 tasks

#Model
1DeepSeek V4.1 Flashhosteddeepseek · 215 tok/s · openrouter
2Qwen3.8 Flashhostedqwen · 64 tok/s · openrouter
3DeepSeek V4 Flash 0731hosteddeepseek · 65 tok/s · openrouter
4Qwen3.8 27Bhostedopenai · 45 tok/s · openai
5Nemotron 3 Ultrahostednvidia · 21 tok/s · openrouter
6Dots3-Note Previewhosteddots-studio · 75 tok/s · openrouter
7Ling-3.0 flashhostedinclusionai · 264 tok/s · openrouter
8Qwen3.5 4B2/9 tasksQ4_K_XL · 26.9 GB · 85 tok/s
9Nanbeige4.2 3B2/9 tasksQ8_0 · 39.5 GB · 33 tok/s
10Qwen3.5 9B2/9 tasksQ8_K_XL · 32.7 GB · 24 tok/s
11Laguna S 2.1hostedpoolside · 48 tok/s · openrouter
12Gemma 4 12B2/9 tasksQ4_K_XL · 26.1 GB · 58 tok/s
13Ling 3.0 Tinyhostedinclusionai · 195 tok/s · openrouter
14LFM2.5 2.6B2/9 tasksQ8_0 · 13.8 GB · 84 tok/s

TOOL Tool Calling·STRUCT Structured Output·RAG RAG / Retrieval QA·CTX Context Recall·CODE Coding·REASON Reasoning & Math·INSTR Instruction Following·CLASS Classification·SUM Summarization

One row per model, ranked by its best build — expand a row for every quantization. Or put two head to head →