config
deployment
unsloth · Q4_K_XL
- quantization
- Q4_K_XL
- harness
- llama.cpp
- context
- 262.1K
- memory
- 26.9 GB
- throughput
- 85 tok/s
- cost / run
- $1.7109est.
- provider
- llama.cpp
- run
- 2026-08-05
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-4B-MTP-GGUF:UD-Q4_K_XL -c 262144Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.
record
per-test results