config
deployment
bartowski · Q8_0
- quantization
- Q8_0
- harness
- llama.cpp
- context
- 193.8K
- memory
- 39.5 GB
- throughput
- 33 tok/s
- cost / run
- $1.5122est.
- provider
- llama.cpp
- run
- 2026-07-30
run this buildllama.cpp
llama-server -hf bartowski/Nanbeige_Nanbeige4.2-3B-GGUF:Q8_0 -c 193792Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.
record
per-test results