# Evaluation cookbooks

Runnable TypeScript recipes that measure whether an agent deployed on Salambo does its job.

The cookbooks are nine small TypeScript programs that evaluate an agent you have deployed. Each one starts Runs through the [TypeScript SDK](/docs/api/runs), reads what happened, and prints a result you can act on. They live in `examples/cookbooks` in the Salambo repository, one file per recipe, and each has a page here that explains the idea and how to change it.

They use only documented `@salambo-ai/sdk` methods, and every Run they start names an agent and an exact version.

## The recipes

| Question                                     | Recipe                                                           | Command                 |
| -------------------------------------------- | ---------------------------------------------------------------- | ----------------------- |
| Did the agent complete the task?             | [Task completion](/docs/cookbooks/task-completion)               | `pnpm task-completion`  |
| How far did it get?                          | [Milestones](/docs/cookbooks/milestones)                         | `pnpm milestones`       |
| Did it stay inside its constraints?          | [Policy checks](/docs/cookbooks/policy-checks)                   | `pnpm policy-checks`    |
| How often does it get it right?              | [Reliability](/docs/cookbooks/reliability)                       | `pnpm reliability`      |
| Is this version better than that one?        | [Compare versions](/docs/cookbooks/compare-versions)             | `pnpm compare-versions` |
| What does a better result cost?              | [Cost and latency](/docs/cookbooks/cost-latency)                 | `pnpm cost-latency`     |
| Is the answer good, when no string can tell? | [LLM judge](/docs/cookbooks/llm-judge)                           | `pnpm llm-judge`        |
| Which real Runs should become tests?         | [Regression set](/docs/cookbooks/regression-set)                 | `pnpm regression-set`   |
| How do I test an agent that acts?            | [Production and evaluation agents](/docs/cookbooks/prod-vs-eval) | `pnpm prod-vs-eval`     |

## Run one

From the repository root:

```bash
pnpm install
export SALAMBO_API_KEY="sk_live_..."
export SALAMBO_BASE_URL="http://localhost:3000"
cd examples/cookbooks
pnpm task-completion --agent agt_01J... --version 3
```

* `SALAMBO_API_KEY` is a key with the `run:read` and `run:write` scopes. [Create one](/docs/guides/create-api-key) in the app.
* `SALAMBO_BASE_URL` is the address of your stack. Use `http://localhost:3000` for a local one. Leave it unset to use the hosted API.
* `--agent` is the agent's ID and `--version` is the exact version to evaluate: the number a Run of that agent reports as `agent_version`. Both can also come from `SALAMBO_AGENT` and `SALAMBO_AGENT_VERSION`.

A recipe that cannot run says what it needs and exits with code 2 before it starts anything: no key, an address nothing answers on, a key Salambo rejects, an agent or version that does not exist, a version that is still building.

## What every recipe does the same way

* **An agent and an exact version.** A Run started on a version does not activate it, so your users keep the version they have while you test another. Every Turn reports the version that executed it, and a recipe stops rather than score a Turn that ran on another one.
* **The Turn is the record, events are evidence.** What a Turn was asked, answered and how it ended comes from the Turn, which is durable. What happened on the way, such as which tools ran, comes from events, which are a rolling window that expires event by event. A recipe that reads events says how complete they were, and reports a finding it cannot settle as unknown, never as a pass. See [Read turn history](/docs/api/runs#read-turn-history) and [History is a rolling window](/docs/api/runs#history-is-a-rolling-window).
* **A limit on Runs in flight.** Every Run starts a sandbox. `--concurrency` (default 3, at most 10) caps how many are in flight, and `--timeout-minutes` (default 15) cancels a Turn that does not finish, and reports it as timed out.
* **No single score.** Each recipe reports per case, or per version, and says what each figure is. A mean hides the case that fails every time. What a second or a token is worth is yours to decide.
* **A result you can gate on.** Exit code 0 means the recipe finished and its checks held (a recipe that only measures has none). 1 means a check failed. 2 means something it needs is missing. 3 means nothing failed but a check could not be settled, because the events it needed were incomplete: a gate must not read that as a pass.

## Your own cases

The built-in cases only need an agent that can answer a question, so every recipe runs on a fresh agent. They say nothing about your agent's real work. Pass your own with `--cases`:

```json
{
  "cases": [
    {
      "id": "refund-window",
      "input": "Can I return an order after 45 days?",
      "mustInclude": ["30 days"],
      "mustNotInclude": ["yes, you can"],
      "rubric": "Passes if the answer states the 30 day return window and offers no exception."
    }
  ]
}
```

`mustInclude` and `mustNotInclude` are checked against the final answer in any letter case. `rubric` is for the [LLM judge](/docs/cookbooks/llm-judge). The [regression set](/docs/cookbooks/regression-set) recipe writes a file in this shape from real Runs.

A recipe that scores answers refuses a case with neither check, because a Turn that completed is not an answer that is right, and counting it as a pass would turn one into the other. Add a check, or pass `--completion-only` to measure only whether Turns complete: the recipe then says that a pass means no more than that.

## Common flags

| Flag                   | Meaning                                                                                                                                                             |
| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| `--agent`, `--version` | The agent and the exact version to evaluate                                                                                                                         |
| `--cases`              | A JSON file of cases instead of the built-in ones                                                                                                                   |
| `--completion-only`    | Task completion, reliability, compare versions and cost and latency only: do not grade answers: accept cases with no check, and measure only whether Turns complete |
| `--concurrency`        | Runs in flight at once. Default 3, at most 10                                                                                                                       |
| `--timeout-minutes`    | How long one Turn may take before it is cancelled. Default 15                                                                                                       |
| `--poll-interval-ms`   | How long a recipe waits before it asks a second time whether a Turn finished. Each later wait is half as long again, up to 10 seconds. Default 1000                 |
| `--json`               | Write the full results to a file. Recipes that score answers record whether answers were graded or only completion was                                              |
| `--help`               | Print the recipe's usage                                                                                                                                            |

## What a recipe cannot know

If the connection fails while a Run is being started and the SDK's own retries do not settle it, the recipe cannot tell whether the request was admitted. It stops and says so, with the idempotency key the start used: a Run may exist and still be working, so look for it in the agent's recent Runs in Salambo. The SDK reuses that key for its own retries, so a lost answer never starts a second Run.

Ctrl-C stops the recipe and cancels the Turns it has started. It does not cut off a start already on its way, because a Run whose start was cut off has no ID to cancel.

## Run tags

A tag is a plain string on a Run, such as `eval` or `regression-candidate`. Every recipe starts its Runs with two tags, `eval` and `cookbook-` followed by the recipe's name, so evaluation traffic can be told from production. The [regression set](/docs/cookbooks/regression-set) recipe finds the Runs a person tagged, with `--tag`. The two calls that use tags are in one file, `src/lib/tags.ts`. See [Tag a run](/docs/api/runs#tag-a-run).

## Two agents

Anything that checks an agent that acts, such as one that refunds orders or sends email, needs a second agent built from the same source and wired to fake services. The [production and evaluation agents](/docs/cookbooks/prod-vs-eval) recipe explains the pattern and runs it.
