# Reliability with pass@k

Run each case several times on one exact version and report pass@1, pass@k and whether every attempt passed.

An agent that passes once may pass one time in three. This recipe runs every case several times, each attempt as a fresh Run with nothing carried over, and reports a rate per case instead of a single pass or fail.

## Run it

```bash
cd examples/cookbooks
pnpm reliability --agent agt_01J... --version 3 --attempts 5 --k 3
```

See [Evaluation cookbooks](/docs/cookbooks/overview) for the key, the address of your stack and the flags every recipe takes.

| Flag                | Meaning                                                                  |
| ------------------- | ------------------------------------------------------------------------ |
| `--attempts`        | Attempts per case. Default 5                                             |
| `--k`               | The `k` in pass\@k. Default 3, at most `--attempts`                      |
| `--require`         | Exit 1 when a case's pass\@k is below this, between 0 and 1              |
| `--completion-only` | Accept cases with no check and measure only whether their Turns complete |
| `--completion-only` | Accept cases with no check and measure only whether their Turns complete |

## What it reports

For a case attempted `n` times with `c` passes:

| Figure       | Meaning                                        | Who it describes                                 |
| ------------ | ---------------------------------------------- | ------------------------------------------------ |
| `pass@1`     | `c / n`, the chance one attempt passes         | A user who gets one try                          |
| `pass@k`     | The chance at least one of `k` attempts passes | A user who may retry `k` times                   |
| `all passed` | Every attempt passed                           | A user who needs it right first time, every time |

pass\@k is the unbiased estimator, `1 - C(n - c, k) / C(n, k)`: it uses all `n` attempts, not only the first `k`, so it does not depend on which attempts happened to come first. With `n = 5`, `c = 2` and `k = 2` it is 0.7. With `k = 1` it is just `c / n`.

```text
case        passed  wrong  errored  pass@1  pass@3  all passed
----------  ------  -----  -------  ------  ------  ----------
arithmetic  3/5     2      0        0.60    1.00    no
json-shape  0/5     4      1        0.00    0.00    no
summary     5/5     0      0        1.00    1.00    yes
```

## Wrong is not errored

An attempt that ends `completed` with a bad answer is `wrong`: a behavior to fix. An attempt whose Turn did not complete, because it failed, stopped part way, was cancelled or ran out of time, is `errored`. Both count as not passing, and both are shown, because they point at different things. A rate with errored attempts in it partly measures the platform, the provider or a tool, so read those Runs before you trust it.

## No average

The recipe prints no mean across cases. A mean hides the case that fails every time behind ones that never do. If you need a gate, `--require 0.8` exits 1 when any single case has a pass\@k under 0.8, and names it.

## How it works

Attempts are independent Runs, so they can run side by side within the concurrency limit. Each one is a separate `runs.execute`, so each gets its own idempotency key. Do not give attempts of one case the same key: a repeated call with the same key returns the same Turn, and the second attempt would replay the first. See [Idempotency](/docs/api/runs#idempotency).

```ts
const observed = await mapWithLimit(
  planRuns(cases, [target], attempts),
  limits.concurrency,
  ({ evalCase, attempt }) => observer.observe(target, evalCase, attempt),
);
```

`planRuns` is in `src/lib/attempts.ts`: every attempt of every case on every target. `mapWithLimit` is in `src/lib/concurrency.ts`. It keeps at most `--concurrency` Runs in flight and returns results in order.

## Change it

More attempts make the rate steadier and cost more Runs. Five attempts distinguish "usually" from "rarely". They cannot distinguish 0.9 from 0.95. Put the cases whose pass\@k is neither 0 nor 1 through more attempts, and start with the [compare versions](/docs/cookbooks/compare-versions) recipe when the question is whether a change helped.
