Laguna S 2.1

poolside/laguna-s-2.1:free

← all builds & scores for Laguna S 2.1
pass rate
27%
64 of 238 tests fully correct
run at a glance
avg throughput
36.1
tok/s · 48 over window
avg output
2.5K
tok / test
avg reasoning
1.9K
tok / test · 76% of output
vram peak
not measured
total tokens
2.1M
1.5M in · 583.6K out
performance

Speed, output & result

Throughputpeak 64 tok/s
7550250
Output
10K7.5K5K2.5K0
result64 pass · 174 fail
tools-ascii-checksum · failtools-isbn-check · passtools-modexp-rotation · failtools-base64-envelope · failtools-log-injection · passtools-threshold-fires · passtools-refund-prorated · failtools-threshold-375gib · passtools-cron-step · passtools-mirror-config · failtools-three-way-split · failtools-window-cross-midnight · passtools-split-fleet · passtools-refund-batch · failtools-fx-repricing · passtools-threshold-matrix · failtools-ratelimit-fleet · failtools-portfolio-rebalance · passtools-maintenance-cascade · passtools-tier-quotes · passtools-fx-normalize · passtools-topo-schedule · failtools-adversarial-dag-dispatch · failtools-adversarial-ledger-netting · failtools-adversarial-calendar-roll · failtools-adversarial-xor-recovery · failtools-adversarial-call-auction · passstruct-string-forensics · passstruct-base-radix · failstruct-collatz-703 · failstruct-mirror-sort · failstruct-utc-forensics · failstruct-typed-traps · passstruct-escape-gauntlet · passstruct-cents-invoice · passstruct-sorted-events · passstruct-derived-consts · passstruct-letter-histogram · failstruct-unit-ladder · passstruct-ledger-aggregates · failstruct-warehouse-census · failstruct-survey-tally · failstruct-timeseries-stats · failstruct-grade-rollup · failstruct-loglevel-census · failstruct-invoice-cascade · failstruct-matrix-margins · failstruct-bigint-chain · passstruct-edge-degrees · failstruct-adversarial-journal-replay · failstruct-adversarial-scc-census · failstruct-adversarial-business-accrual · failstruct-adversarial-stack-transcript · failstruct-adversarial-generated-matrix · failrag-fx-ledger · failrag-roster-window · failrag-quota-mixed-units · passrag-escalation-parity · failrag-derived-retention · failrag-failed-minutes · failrag-weighted-sla · failrag-sku-delta · failrag-version-pin · passrag-grep-count · failrag-two-phase-quota · failrag-abstain-inference-bait · passrag-weighted-sla-fleet · failrag-fx-ledger-balance · failrag-retention-audit · failrag-inventory-final · failrag-failed-run-minutes · failrag-ticket-filter · failrag-quota-allocation · passrag-invoice-aging · failrag-headcount-dept · passrag-weighted-score · failrag-adversarial-bitemporal-policy · failrag-adversarial-lineage-incident-join · failrag-adversarial-contract-precedence · failrag-adversarial-access-resolution · failrag-adversarial-provenance-corrections · failctx-len-10 · passctx-len-25 · passctx-len-50 · passctx-len-75 · passctx-len-100 · passctx-recency-override · failctx-hard-negatives · failctx-multi-needle-sum · passctx-count-exact-tag · failctx-multihop-chain · passctx-filter-threshold · failctx-conflict-scope · passctx-nth-occurrence · failctx-adversarial-pointer-cycle · failctx-adversarial-seventh-checkpoint · failctx-adversarial-scoped-revisions · failctx-adversarial-multihop-appendices · failctx-adversarial-conjunctive-needle · passcode-stack-vm · passcode-glob-match · failcode-iso-weekdate · failcode-crc32-utf8 · failcode-introot-dec · passcode-reverse-clusters · failcode-bankers-round-dec · failcode-iso-duration · failcode-topo-or-cycles · failcode-decode-nested · failcode-normalize-ranges · passcode-justify-text · passcode-glyph-fold · failcode-toll-fare · passcode-rune-score · passcode-fold-sum · passcode-fib-sum-mod · failcode-big-modpow · failcode-grid-paths · passcode-mod-luhn · failcode-token-stack-vm · failcode-clock-chain · failcode-adversarial-huge-recurrence · failcode-adversarial-assignment-lexicographic · failcode-adversarial-regex-language-count · failcode-adversarial-periodic-shortest-path · failcode-adversarial-exact-rational-parser · passreason-mult-15x15 · failreason-digit-sum-2-333 · failreason-modexp · failreason-collatz-steps · failreason-smallest-multiple · failreason-prime-census · failreason-anchored-sum · failreason-forgetful-host · passreason-wide-boat · failreason-long-multiplication · failreason-digit-sum-power · failreason-century-leap · passreason-crt-trio · passreason-hexad-chain · failreason-quad-product · failreason-power-tower · failreason-recurrence-sum · failreason-crt-fib-inverse · failreason-base-digit-chain · failreason-triple-modexp · failreason-factorial-sum · failreason-fib-pair-digits · failreason-grand-octet · failreason-adversarial-recurrence-sum · failreason-adversarial-exactly-three-divisors · failreason-adversarial-noncoprime-crt · failreason-adversarial-million-grid · failreason-adversarial-combinatorial-fusion · failinstr-mirror-90 · failinstr-reverse-lex-30 · failinstr-rot13-swapcase · failinstr-every-third · failinstr-acrostic-robust · failinstr-e-atlas · failinstr-acrostic-solid · failinstr-reverse-long-word · failinstr-abc-ladder · failinstr-o-census · failinstr-word-lathe · failinstr-self-count · passinstr-interleave · failinstr-word-pipeline · failinstr-position-cipher · failinstr-interleave-prune · failinstr-vigenere · failinstr-columnar-transpose · failinstr-rle-encode · failinstr-reverse-vowelcase · failinstr-atbash-shift · failinstr-caesar-decode · failinstr-mega-chain · failinstr-adversarial-coprime-block-cipher · failinstr-adversarial-token-permutation · failinstr-adversarial-grid-selection · failinstr-adversarial-bwt-move-to-front · failinstr-adversarial-byte-pipeline · failclass-bracket-mode-a · passclass-bracket-mode-b · failclass-popcount-mod4 · failclass-s-census · failclass-factor-trio-a · failclass-factor-trio-b · failclass-relation-triple · failclass-anagram-triple · failclass-weekday-far · failclass-needle-census · passclass-empty-set-syllogism · passclass-mod-residue · failclass-char-atlas · failclass-longest-run · passclass-weighted-checksum · failclass-cipher-runsum · failclass-mod97-residue · failclass-rolling-hash · failclass-pattern-census · failclass-bracket-maxdepth · failclass-weighted-digit-mod · failclass-ascending-transitions · failclass-frequency-signature · failclass-case-swings · failclass-adversarial-planted-3sat · failclass-adversarial-dfa-equivalence · failclass-adversarial-graph-family · failclass-adversarial-lfsr-signature · failclass-adversarial-simpson-direction · passsum-log-forensics · failsum-outage-union · failsum-dependency-diff · failsum-session-census · failsum-expense-exact · failsum-incident-json · passsum-budgeted-facts · passsum-spelled-out · passsum-clean-headline · failsum-forced-paraphrase · passsum-priority-actions · passsum-growth-delta · passsum-reimbursement-total · failsum-outage-window-union · failsum-lockfile-diff · failsum-event-stream-census · failsum-log-severity · failsum-poll-tally · failsum-expense-oneliner · failsum-headcount-24mo · failsum-incident-timeline · passsum-revenue-growth · passsum-adversarial-release-replay · failsum-adversarial-outage-concurrency · failsum-adversarial-trial-meta · passsum-adversarial-consolidation · passsum-adversarial-decision-log · pass
TOOL
STRUCT
RAG
CTX
CODE
REASON
INSTR
CLASS
SUM
memorynot measured

one reading per test, in suite order · the result tape marks each test pass (dim) or fail (bright).

by category

Pass rate per category

9
Tool Calling13/2748%score 60
Structured Output8/2730%score 30
RAG / Retrieval QA5/2719%score 19
Context Recall9/1850%score 50
Coding9/2733%score 54
Reasoning & Math3/2811%score 11
Instruction Following1/284%score 4
Classification5/2917%score 17
Summarization11/2741%score 41
responses

Test-by-test results

Tool Calling13/2748%score 60
Structured Output8/2730%score 30
RAG / Retrieval QA5/2719%score 19
Context Recall9/1850%score 50
Coding9/2733%score 54
Reasoning & Math3/2811%score 11
Instruction Following1/284%score 4
Classification5/2917%score 17
Summarization11/2741%score 41
run record
run

This run

suite
starter@v1
ran at
2026-08-06 05:20 UTC
record id 
config

Deployment configuration

harness
openrouter
quantization
context
262.1K
temperature
0
multi-token prediction
engine
v1.0.0
provider
openrouter
category call budgetreasoningtotal
Tool Calling4.1K/ high8.2K/ call
Structured Output2.0K/ medium4.1K/ call
RAG / Retrieval QA2.0K/ medium4.1K/ call
Context Recall2.0K/ medium4.1K/ call
Coding4.1K/ high8.2K/ call
Reasoning & Math4.1K/ high8.2K/ call
Instruction Following2.0K/ medium4.1K/ call
Classification2.0K/ medium4.1K/ call
Summarization2.0K/ medium4.1K/ call
run this buildOpenRouterhosted
curl https://openrouter.ai/api/v1/chat/completions \
  -H "Authorization: Bearer $OPENROUTER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model":"poolside/laguna-s-2.1:free","messages":[{"role":"user","content":"Hello"}]}'

Hosted API — not self-hostable. This is the call the benchmark made; try it live in the Playground.

measured

How it was measured

latency
12275636 ms
cost
$0
memory samples

tok/s and latency measure the scored generation window; per-test figures above are the readings for each case. Memory (VRAM, unified RAM, or host RAM) is the peak footprint polled throughout the run, or blank when no probe could attribute it — never a hand-typed number.

recipe
prompts → hashes

Reproduce this run

The private prompts never leave the suite. What’s public is the command, the pinned suite, and the hashes that stand in for the tests — enough to re-run the exact config against your own copy of the suite.

command
npm run engine -- run --provider openrouter --model poolside/laguna-s-2.1:free --suite suites/starter.suite.json --name 'Laguna S 2.1'
suite hash
config hash
suite pinned to
starter@v1