Nanbeige4.2 3B

bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0

llama.cpp1 build benchmarkedbest build bartowski · Q8_039.5 GB33 tok/s
best build

Scores at a glance

measured avg36
Tool Calling476/27
Reasoning & Math257/28
0255075100

0–100 per task · n/m = cases passed (scores credit partial passes) · measured average covers 2/9 suite tasks

best build
config
deployment

bartowski · Q8_0

quantization
Q8_0
harness
llama.cpp
context
193.8K
memory
39.5 GB
throughput
33 tok/s
cost / run
$1.5122est.
provider
llama.cpp
run
2026-07-30
run this buildllama.cpp
llama-server -hf bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0 -c 193792

Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.

record
per-test results

Full run record

Per-test scores, full transcripts, and the reproduce recipe.

open the run record →