Gemma 4 12B

unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL

← all builds & scores for Gemma 4 12B
pass rate
16%
9 of 55 tests fully correct
run at a glance
avg throughput
58.2
tok/s · 58 over window
avg output
4.7K
tok / test
avg reasoning
3.7K
tok / test · 78% of output
unified ram
26.1 GB
peak memory
total tokens
427.6K
167.8K in · 259.7K out
performance

Speed, output & result

Throughputpeak 68 tok/s
7550250
Output
12.5K10K7.5K5K2.5K0
result9 pass · 46 fail
tools-ascii-checksum · failtools-isbn-check · failtools-modexp-rotation · failtools-base64-envelope · failtools-log-injection · passtools-threshold-fires · failtools-refund-prorated · failtools-threshold-375gib · failtools-cron-step · failtools-mirror-config · failtools-three-way-split · failtools-window-cross-midnight · failtools-split-fleet · failtools-refund-batch · failtools-fx-repricing · failtools-threshold-matrix · failtools-ratelimit-fleet · passtools-portfolio-rebalance · failtools-maintenance-cascade · failtools-tier-quotes · failtools-fx-normalize · failtools-topo-schedule · failtools-adversarial-dag-dispatch · failtools-adversarial-ledger-netting · failtools-adversarial-calendar-roll · failtools-adversarial-xor-recovery · failtools-adversarial-call-auction · failreason-mult-15x15 · failreason-digit-sum-2-333 · failreason-modexp · failreason-collatz-steps · passreason-smallest-multiple · failreason-prime-census · failreason-anchored-sum · passreason-forgetful-host · passreason-wide-boat · passreason-long-multiplication · passreason-digit-sum-power · failreason-century-leap · passreason-crt-trio · passreason-hexad-chain · failreason-quad-product · failreason-power-tower · failreason-recurrence-sum · failreason-crt-fib-inverse · failreason-base-digit-chain · failreason-triple-modexp · failreason-factorial-sum · failreason-fib-pair-digits · failreason-grand-octet · failreason-adversarial-recurrence-sum · failreason-adversarial-exactly-three-divisors · failreason-adversarial-noncoprime-crt · failreason-adversarial-million-grid · failreason-adversarial-combinatorial-fusion · fail
TOOL
REASON
memory25.8 GB peak

one reading per test, in suite order · the result tape marks each test pass (dim) or fail (bright).

by category

Pass rate per category

2
Tool Calling2/277%score 30
Reasoning & Math7/2825%score 25
responses

Test-by-test results

Tool Calling2/277%score 30
Reasoning & Math7/2825%score 25
run record
run

This run

suite
starter@v1
ran at
2026-07-30 05:46 UTC
record id 
config

Deployment configuration

harness
llama.cpp
quantization
Q4_K_XL
context
262.1K
temperature
0
multi-token prediction
engine
v1.0.0
provider
llama.cpp
category call budgetreasoningtotal
Tool Calling4.1K/ high8.2K/ call
Structured Output2.0K/ medium4.1K/ call
RAG / Retrieval QA2.0K/ medium4.1K/ call
Context Recall2.0K/ medium4.1K/ call
Coding4.1K/ high8.2K/ call
Reasoning & Math4.1K/ high8.2K/ call
Instruction Following2.0K/ medium4.1K/ call
Classification2.0K/ medium4.1K/ call
Summarization2.0K/ medium4.1K/ call
run this buildllama.cpp
llama-server -hf unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL -c 262144

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

measured

How it was measured

latency
4508689 ms
cost
$1.0308est.
memory samples
4559

tok/s and latency measure the scored generation window; per-test figures above are the readings for each case. Memory (VRAM, unified RAM, or host RAM) is the peak footprint polled throughout the run — measured via process rss across 4559 samples, or blank when no probe could attribute it — never a hand-typed number.

recipe
prompts → hashes

Reproduce this run

The private prompts never leave the suite. What’s public is the command, the pinned suite, and the hashes that stand in for the tests — enough to re-run the exact config against your own copy of the suite.

command
npm run engine -- run --provider llama.cpp --model unsloth/gemma-4-12b-it-GGUF:UD-Q4_K_XL --suite suites/starter.suite.json --name 'Gemma 4 12B' --category 'tool-calling,reasoning'
suite hash
config hash
suite pinned to
starter@v1