config
deployment
unsloth · Q8_K_XL
- quantization
- Q8_K_XL
- harness
- llama.cpp
- context
- 262.1K
- memory
- 32.7 GB
- throughput
- 24 tok/s
- cost / run
- $7.2603est.
- provider
- llama.cpp
- run
- 2026-07-30
run this buildllama.cpp
llama-server -hf unsloth/Qwen3.5-9B-GGUF:UD-Q8_K_XL -c 262144Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.
record
per-test results