Browse documentation
Reliability with pass@k
Run each case several times on one exact version and report pass@1, pass@k and whether every attempt passed.
View MarkdownAn agent that passes once may pass one time in three. This recipe runs every case several times, each attempt as a fresh Run with nothing carried over, and reports a rate per case instead of a single pass or fail.
Run it
cd examples/cookbooks
pnpm reliability --agent agt_01J... --version 3 --attempts 5 --k 3See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.
| Flag | Meaning |
|---|---|
--attempts | Attempts per case. Default 5 |
--k | The k in pass@k. Default 3, at most --attempts |
--require | Exit 1 when a case's pass@k is below this, between 0 and 1 |
--completion-only | Accept cases with no check and measure only whether their Turns complete |
--completion-only | Accept cases with no check and measure only whether their Turns complete |
What it reports
For a case attempted n times with c passes:
| Figure | Meaning | Who it describes |
|---|---|---|
pass@1 | c / n, the chance one attempt passes | A user who gets one try |
pass@k | The chance at least one of k attempts passes | A user who may retry k times |
all passed | Every attempt passed | A user who needs it right first time, every time |
pass@k is the unbiased estimator, 1 - C(n - c, k) / C(n, k): it uses all n attempts, not only the first k, so it does not depend on which attempts happened to come first. With n = 5, c = 2 and k = 2 it is 0.7. With k = 1 it is just c / n.
case passed wrong errored pass@1 pass@3 all passed
---------- ------ ----- ------- ------ ------ ----------
arithmetic 3/5 2 0 0.60 1.00 no
json-shape 0/5 4 1 0.00 0.00 no
summary 5/5 0 0 1.00 1.00 yesWrong is not errored
An attempt that ends completed with a bad answer is wrong: a behavior to fix. An attempt whose Turn did not complete, because it failed, stopped part way, was cancelled or ran out of time, is errored. Both count as not passing, and both are shown, because they point at different things. A rate with errored attempts in it partly measures the platform, the provider or a tool, so read those Runs before you trust it.
No average
The recipe prints no mean across cases. A mean hides the case that fails every time behind ones that never do. If you need a gate, --require 0.8 exits 1 when any single case has a pass@k under 0.8, and names it.
How it works
Attempts are independent Runs, so they can run side by side within the concurrency limit. Each one is a separate runs.execute, so each gets its own idempotency key. Do not give attempts of one case the same key: a repeated call with the same key returns the same Turn, and the second attempt would replay the first. See Idempotency.
const observed = await mapWithLimit(
planRuns(cases, [target], attempts),
limits.concurrency,
({ evalCase, attempt }) => observer.observe(target, evalCase, attempt),
);planRuns is in src/lib/attempts.ts: every attempt of every case on every target. mapWithLimit is in src/lib/concurrency.ts. It keeps at most --concurrency Runs in flight and returns results in order.
Change it
More attempts make the rate steadier and cost more Runs. Five attempts distinguish "usually" from "rarely". They cannot distinguish 0.9 from 0.95. Put the cases whose pass@k is neither 0 nor 1 through more attempts, and start with the compare versions recipe when the question is whether a change helped.