Gemma 4 12B

unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL

llama.cpp1 build benchmarkedbest build unsloth · Q4_K_XL26.1 GB58 tok/s
best build

Scores at a glance

measured avg28
Tool Calling302/27
Reasoning & Math257/28
0255075100

0–100 per task · n/m = cases passed (scores credit partial passes) · measured average covers 2/9 suite tasks

best build
config
deployment

unsloth · Q4_K_XL

quantization
Q4_K_XL
harness
llama.cpp
context
262.1K
memory
26.1 GB
throughput
58 tok/s
cost / run
$1.0308est.
provider
llama.cpp
run
2026-07-30
run this buildllama.cpp
llama-server -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

record
per-test results

Full run record

Per-test scores, full transcripts, and the reproduce recipe.

open the run record →