Qwen3.5 4B

unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL

← all builds & scores for Qwen3.5 4B
pass rate
24%
13 of 55 tests fully correct
run at a glance
avg throughput
88.3
tok/s · 85 over window
avg output
10.6K
tok / test
avg reasoning
8.5K
tok / test · 80% of output
unified ram
26.9 GB
peak memory
total tokens
1.0M
465.7K in · 583.9K out
performance

Speed, output & result

Throughputpeak 106 tok/s
1251007550250
Output
120K100K80K60K40K20K0
result13 pass · 42 fail
tools-ascii-checksum · failtools-isbn-check · passtools-modexp-rotation · failtools-base64-envelope · failtools-log-injection · passtools-threshold-fires · passtools-refund-prorated · passtools-threshold-375gib · failtools-cron-step · failtools-mirror-config · failtools-three-way-split · passtools-window-cross-midnight · failtools-split-fleet · failtools-refund-batch · failtools-fx-repricing · failtools-threshold-matrix · failtools-ratelimit-fleet · passtools-portfolio-rebalance · failtools-maintenance-cascade · failtools-tier-quotes · passtools-fx-normalize · passtools-topo-schedule · failtools-adversarial-dag-dispatch · failtools-adversarial-ledger-netting · failtools-adversarial-calendar-roll · failtools-adversarial-xor-recovery · failtools-adversarial-call-auction · failreason-mult-15x15 · failreason-digit-sum-2-333 · failreason-modexp · failreason-collatz-steps · passreason-smallest-multiple · failreason-prime-census · failreason-anchored-sum · passreason-forgetful-host · passreason-wide-boat · passreason-long-multiplication · failreason-digit-sum-power · failreason-century-leap · passreason-crt-trio · failreason-hexad-chain · failreason-quad-product · failreason-power-tower · failreason-recurrence-sum · failreason-crt-fib-inverse · failreason-base-digit-chain · failreason-triple-modexp · failreason-factorial-sum · failreason-fib-pair-digits · failreason-grand-octet · failreason-adversarial-recurrence-sum · failreason-adversarial-exactly-three-divisors · failreason-adversarial-noncoprime-crt · failreason-adversarial-million-grid · failreason-adversarial-combinatorial-fusion · fail
TOOL
REASON
memory26.0 GB peak

one reading per test, in suite order · the result tape marks each test pass (dim) or fail (bright).

by category

Pass rate per category

2
Tool Calling8/2730%score 55
Reasoning & Math5/2818%score 18
responses

Test-by-test results

Tool Calling8/2730%score 55
Reasoning & Math5/2818%score 18
run record
run

This run

suite
starter@v1
ran at
2026-08-05 14:15 UTC
record id 
config

Deployment configuration

harness
llama.cpp
quantization
Q4_K_XL
context
262.1K
temperature
0
multi-token prediction
engine
v1.0.0
provider
llama.cpp
category call budgetreasoningtotal
Tool Calling4.1K/ high8.2K/ call
Structured Output2.0K/ medium4.1K/ call
RAG / Retrieval QA2.0K/ medium4.1K/ call
Context Recall2.0K/ medium4.1K/ call
Coding4.1K/ high8.2K/ call
Reasoning & Math4.1K/ high8.2K/ call
Instruction Following2.0K/ medium4.1K/ call
Classification2.0K/ medium4.1K/ call
Summarization2.0K/ medium4.1K/ call
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

measured

How it was measured

latency
6853321 ms
cost
$1.7109est.
memory samples
6901

tok/s and latency measure the scored generation window; per-test figures above are the readings for each case. Memory (VRAM, unified RAM, or host RAM) is the peak footprint polled throughout the run — measured via process rss across 6901 samples, or blank when no probe could attribute it — never a hand-typed number.

recipe
prompts → hashes

Reproduce this run

The private prompts never leave the suite. What’s public is the command, the pinned suite, and the hashes that stand in for the tests — enough to re-run the exact config against your own copy of the suite.

command
npm run engine -- run --provider llama.cpp --model unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL --suite suites/starter.suite.json --name 'Qwen3.5 4B' --category 'tool-calling,reasoning'
suite hash
config hash
suite pinned to
starter@v1