Qwen3.5 9B

unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL

← all builds & scores for Qwen3.5 9B
pass rate
24%
13 of 55 tests fully correct
run at a glance
avg throughput
24.1
tok/s · 24 over window
avg output
12.4K
tok / test
avg reasoning
10.9K
tok / test · 88% of output
unified ram
32.7 GB
peak memory
total tokens
1.2M
559.2K in · 683.4K out
performance

Speed, output & result

Throughputpeak 24 tok/s
250
Output
120K100K80K60K40K20K0
result13 pass · 42 fail
tools-ascii-checksum · failtools-isbn-check · passtools-modexp-rotation · failtools-base64-envelope · failtools-log-injection · failtools-threshold-fires · passtools-refund-prorated · passtools-threshold-375gib · failtools-cron-step · failtools-mirror-config · failtools-three-way-split · passtools-window-cross-midnight · failtools-split-fleet · failtools-refund-batch · failtools-fx-repricing · failtools-threshold-matrix · passtools-ratelimit-fleet · failtools-portfolio-rebalance · failtools-maintenance-cascade · failtools-tier-quotes · passtools-fx-normalize · failtools-topo-schedule · passtools-adversarial-dag-dispatch · failtools-adversarial-ledger-netting · failtools-adversarial-calendar-roll · failtools-adversarial-xor-recovery · failtools-adversarial-call-auction · failreason-mult-15x15 · failreason-digit-sum-2-333 · failreason-modexp · failreason-collatz-steps · failreason-smallest-multiple · failreason-prime-census · failreason-anchored-sum · passreason-forgetful-host · passreason-wide-boat · passreason-long-multiplication · failreason-digit-sum-power · passreason-century-leap · passreason-crt-trio · passreason-hexad-chain · failreason-quad-product · failreason-power-tower · failreason-recurrence-sum · failreason-crt-fib-inverse · failreason-base-digit-chain · failreason-triple-modexp · failreason-factorial-sum · failreason-fib-pair-digits · failreason-grand-octet · failreason-adversarial-recurrence-sum · failreason-adversarial-exactly-three-divisors · failreason-adversarial-noncoprime-crt · failreason-adversarial-million-grid · failreason-adversarial-combinatorial-fusion · fail
TOOL
REASON
memory32.5 GB peak

one reading per test, in suite order · the result tape marks each test pass (dim) or fail (bright).

by category

Pass rate per category

2
Tool Calling7/2726%score 49
Reasoning & Math6/2821%score 21
responses

Test-by-test results

Tool Calling7/2726%score 49
Reasoning & Math6/2821%score 21
run record
run

This run

suite
starter@v1
ran at
2026-07-30 18:51 UTC
record id 
config

Deployment configuration

harness
llama.cpp
quantization
Q8_K_XL
context
262.1K
temperature
0
multi-token prediction
engine
v1.0.0
provider
llama.cpp
category call budgetreasoningtotal
Tool Calling4.1K/ high8.2K/ call
Structured Output2.0K/ medium4.1K/ call
RAG / Retrieval QA2.0K/ medium4.1K/ call
Context Recall2.0K/ medium4.1K/ call
Coding4.1K/ high8.2K/ call
Reasoning & Math4.1K/ high8.2K/ call
Instruction Following2.0K/ medium4.1K/ call
Classification2.0K/ medium4.1K/ call
Summarization2.0K/ medium4.1K/ call
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

measured

How it was measured

latency
28752657 ms
cost
$7.2603est.
memory samples
28763

tok/s and latency measure the scored generation window; per-test figures above are the readings for each case. Memory (VRAM, unified RAM, or host RAM) is the peak footprint polled throughout the run — measured via process rss across 28763 samples, or blank when no probe could attribute it — never a hand-typed number.

recipe
prompts → hashes

Reproduce this run

The private prompts never leave the suite. What’s public is the command, the pinned suite, and the hashes that stand in for the tests — enough to re-run the exact config against your own copy of the suite.

command
npm run engine -- run --provider llama.cpp --model unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL --suite suites/starter.suite.json --name 'Qwen3.5 9B' --category 'tool-calling,reasoning'
suite hash
config hash
suite pinned to
starter@v1