config
deployment
liquidai · Q8_0
- quantization
- Q8_0
- harness
- llama.cpp
- context
- 128K
- memory
- 13.8 GB
- throughput
- 84 tok/s
- cost / run
- $1.1119est.
- provider
- llama.cpp
- run
- 2026-08-05
run this buildllama.cpp
llama-server -hf LiquidAI/LFM2.5-2.6B-GGUF:Q8_0 -c 128000Pulls the GGUF from Hugging Face and serves an OpenAI-compatible endpoint on :8080.
record
per-test results