config
deployment
unsloth · Q4_K_XL
- quantization
- Q4_K_XL
- harness
- llama.cpp
- context
- 262.1K
- memory
- 26.1 GB
- throughput
- 58 tok/s
- cost / run
- $1.0308est.
- provider
- llama.cpp
- run
- 2026-07-30
run this buildllama.cpp
llama-server -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL -c 262144Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.
record
per-test results