Which open-weight model should you run?
Real build-tasks, ranked for the hardware you have. The hosted frontier is the reference line — every score reproducible.
the headline matchupQwen3.5 4B 37 vs DeepSeek V4.1 Flash 60
the full matchup →14 models·238 cases·9 tasks
#Model
TOOL Tool Calling·STRUCT Structured Output·RAG RAG / Retrieval QA·CTX Context Recall·CODE Coding·REASON Reasoning & Math·INSTR Instruction Following·CLASS Classification·SUM Summarization
One row per model, ranked by its best build — expand a row for every quantization. Or put two head to head →