Qwen3.5 4B

unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL

llama.cpp1 build benchmarkedbest build unsloth · Q4_K_XL26.9 GB85 tok/s
best build

Scores at a glance

measured avg37
Tool Calling558/27
Reasoning & Math185/28
0255075100

0–100 per task · n/m = cases passed (scores credit partial passes) · measured average covers 2/9 suite tasks

best build
config
deployment

unsloth · Q4_K_XL

quantization
Q4_K_XL
harness
llama.cpp
context
262.1K
memory
26.9 GB
throughput
85 tok/s
cost / run
$1.7109est.
provider
llama.cpp
run
2026-08-05
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

record
per-test results

Full run record

Per-test scores, full transcripts, and the reproduce recipe.

open the run record →