Nanbeige4.2 3B

bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0

← all builds & scores for Nanbeige4.2 3B
pass rate
24%
13 of 55 tests fully correct
run at a glance
avg throughput
33.7
tok/s · 33 over window
avg output
5.4K
tok / test
avg reasoning
3.6K
tok / test · 66% of output
unified ram
39.5 GB
peak memory
total tokens
360.4K
65.2K in · 295.2K out
performance

Speed, output & result

Throughputpeak 36 tok/s
50250
Output
15K12.5K10K7.5K5K2.5K0
result13 pass · 42 fail
tools-ascii-checksum · failtools-isbn-check · failtools-modexp-rotation · failtools-base64-envelope · failtools-log-injection · passtools-threshold-fires · passtools-refund-prorated · failtools-threshold-375gib · failtools-cron-step · failtools-mirror-config · failtools-three-way-split · failtools-window-cross-midnight · failtools-split-fleet · failtools-refund-batch · passtools-fx-repricing · failtools-threshold-matrix · failtools-ratelimit-fleet · passtools-portfolio-rebalance · failtools-maintenance-cascade · passtools-tier-quotes · failtools-fx-normalize · passtools-topo-schedule · failtools-adversarial-dag-dispatch · failtools-adversarial-ledger-netting · failtools-adversarial-calendar-roll · failtools-adversarial-xor-recovery · failtools-adversarial-call-auction · failreason-mult-15x15 · failreason-digit-sum-2-333 · failreason-modexp · passreason-collatz-steps · failreason-smallest-multiple · failreason-prime-census · failreason-anchored-sum · passreason-forgetful-host · passreason-wide-boat · failreason-long-multiplication · passreason-digit-sum-power · passreason-century-leap · passreason-crt-trio · passreason-hexad-chain · failreason-quad-product · failreason-power-tower · failreason-recurrence-sum · failreason-crt-fib-inverse · failreason-base-digit-chain · failreason-triple-modexp · failreason-factorial-sum · failreason-fib-pair-digits · failreason-grand-octet · failreason-adversarial-recurrence-sum · failreason-adversarial-exactly-three-divisors · failreason-adversarial-noncoprime-crt · failreason-adversarial-million-grid · failreason-adversarial-combinatorial-fusion · fail
TOOL
REASON
memory39.1 GB peak

one reading per test, in suite order · the result tape marks each test pass (dim) or fail (bright).

by category

Pass rate per category

2
Tool Calling6/2722%score 47
Reasoning & Math7/2825%score 25
responses

Test-by-test results

Tool Calling6/2722%score 47
Reasoning & Math7/2825%score 25
run record
run

This run

suite
starter@v1
ran at
2026-07-30 14:10 UTC
record id 
config

Deployment configuration

harness
llama.cpp
quantization
Q8_0
context
193.8K
temperature
0
multi-token prediction
engine
v1.0.0
provider
llama.cpp
category call budgetreasoningtotal
Tool Calling4.1K/ high8.2K/ call
Structured Output2.0K/ medium4.1K/ call
RAG / Retrieval QA2.0K/ medium4.1K/ call
Context Recall2.0K/ medium4.1K/ call
Coding4.1K/ high8.2K/ call
Reasoning & Math4.1K/ high8.2K/ call
Instruction Following2.0K/ medium4.1K/ call
Classification2.0K/ medium4.1K/ call
Summarization2.0K/ medium4.1K/ call
run this buildllama.cpp
llama-server -hf bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0 -c 193792

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

measured

How it was measured

latency
8919435 ms
cost
$1.5122est.
memory samples
8965

tok/s and latency measure the scored generation window; per-test figures above are the readings for each case. Memory (VRAM, unified RAM, or host RAM) is the peak footprint polled throughout the run — measured via process rss across 8965 samples, or blank when no probe could attribute it — never a hand-typed number.

recipe
prompts → hashes

Reproduce this run

The private prompts never leave the suite. What’s public is the command, the pinned suite, and the hashes that stand in for the tests — enough to re-run the exact config against your own copy of the suite.

command
npm run engine -- run --provider llama.cpp --model bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0 --suite suites/starter.suite.json --name 'Nanbeige4.2 3B' --category 'tool-calling,reasoning'
suite hash
config hash
suite pinned to
starter@v1