Qwen3.5 9B

unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL

llama.cpp1 build benchmarkedbest build unsloth · Q8_K_XL32.7 GB24 tok/s
best build

Scores at a glance

measured avg35
Tool Calling497/27
Reasoning & Math216/28
0255075100

0–100 per task · n/m = cases passed (scores credit partial passes) · measured average covers 2/9 suite tasks

best build
config
deployment

unsloth · Q8_K_XL

quantization
Q8_K_XL
harness
llama.cpp
context
262.1K
memory
32.7 GB
throughput
24 tok/s
cost / run
$7.2603est.
provider
llama.cpp
run
2026-07-30
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

record
per-test results

Full run record

Per-test scores, full transcripts, and the reproduce recipe.

open the run record →