Salambo
Browse documentation
Evaluation cookbooksReliability with pass@k

Reliability with pass@k

Run each case several times on one exact version and report pass@1, pass@k and whether every attempt passed.

View Markdown

An agent that passes once may pass one time in three. This recipe runs every case several times, each attempt as a fresh Run with nothing carried over, and reports a rate per case instead of a single pass or fail.

Run it

bash
cd examples/cookbooks
pnpm reliability --agent agt_01J... --version 3 --attempts 5 --k 3

See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.

FlagMeaning
--attemptsAttempts per case. Default 5
--kThe k in pass@k. Default 3, at most --attempts
--requireExit 1 when a case's pass@k is below this, between 0 and 1
--completion-onlyAccept cases with no check and measure only whether their Turns complete
--completion-onlyAccept cases with no check and measure only whether their Turns complete

What it reports

For a case attempted n times with c passes:

FigureMeaningWho it describes
pass@1c / n, the chance one attempt passesA user who gets one try
pass@kThe chance at least one of k attempts passesA user who may retry k times
all passedEvery attempt passedA user who needs it right first time, every time

pass@k is the unbiased estimator, 1 - C(n - c, k) / C(n, k): it uses all n attempts, not only the first k, so it does not depend on which attempts happened to come first. With n = 5, c = 2 and k = 2 it is 0.7. With k = 1 it is just c / n.

text
case        passed  wrong  errored  pass@1  pass@3  all passed
----------  ------  -----  -------  ------  ------  ----------
arithmetic  3/5     2      0        0.60    1.00    no
json-shape  0/5     4      1        0.00    0.00    no
summary     5/5     0      0        1.00    1.00    yes

Wrong is not errored

An attempt that ends completed with a bad answer is wrong: a behavior to fix. An attempt whose Turn did not complete, because it failed, stopped part way, was cancelled or ran out of time, is errored. Both count as not passing, and both are shown, because they point at different things. A rate with errored attempts in it partly measures the platform, the provider or a tool, so read those Runs before you trust it.

No average

The recipe prints no mean across cases. A mean hides the case that fails every time behind ones that never do. If you need a gate, --require 0.8 exits 1 when any single case has a pass@k under 0.8, and names it.

How it works

Attempts are independent Runs, so they can run side by side within the concurrency limit. Each one is a separate runs.execute, so each gets its own idempotency key. Do not give attempts of one case the same key: a repeated call with the same key returns the same Turn, and the second attempt would replay the first. See Idempotency.

ts
const observed = await mapWithLimit(
  planRuns(cases, [target], attempts),
  limits.concurrency,
  ({ evalCase, attempt }) => observer.observe(target, evalCase, attempt),
);

planRuns is in src/lib/attempts.ts: every attempt of every case on every target. mapWithLimit is in src/lib/concurrency.ts. It keeps at most --concurrency Runs in flight and returns results in order.

Change it

More attempts make the rate steadier and cost more Runs. Five attempts distinguish "usually" from "rarely". They cannot distinguish 0.9 from 0.95. Put the cases whose pass@k is neither 0 nor 1 through more attempts, and start with the compare versions recipe when the question is whether a change helped.