Browse documentation
Regression set from tagged Runs
Collect Runs that went wrong, read what each was asked from its Turn history, write them as cases, and replay them on an exact version.
View MarkdownProduction is where the cases you did not think of turn up. When someone finds a Run that went badly, they tag it with a plain string such as regression-candidate. This recipe collects those Runs, turns each one into a case, writes the cases to a file, and replays them on a version you name. You learn whether a fix fixed it, and whether anything that used to work stopped.
Run it
By Run ID, which works today:
cd examples/cookbooks
pnpm regression-set \
--run-ids run_01J...,run_01K... \
--agent agt_01L... --version 3 \
--out regression-set.jsonBy the most recent finished Runs of an agent, when you have no IDs yet:
pnpm regression-set --recent 20 --source-agent agt_01J... \
--agent agt_01L... --version 3By tag:
pnpm regression-set --tag regression-candidate --source-agent agt_01J... \
--agent agt_01L... --version 3Replay goes to a separate evaluation agent: the recipe refuses to replay a Run's request on the agent that Run came from, because that repeats whatever the agent did the first time, such as a refund or an email. Pass --allow-source-agent only for an agent with no side effects.
Choose the Runs with exactly one of --tag, --run-ids and --recent. --recent takes up to 100 finished Runs of the source agent, including evaluation Runs, so point it at an agent that only serves real traffic. With --no-replay the recipe only writes the file. It never replaces a file that is already there, because that file may hold checks you added by hand: choose another --out, or pass --overwrite, which writes the whole replacement beside the file before moving it into place, so a failed write leaves the old file as it was, and never leaves the replacement more open than the file was. A file that others could read is replaced by one that only its owner can read, and the recipe says so. The file holds what Runs were asked, which can be customer text, so a new one is readable by its owner alone. A symbolic link is refused: pass the path of the file itself. The --json results file is owner-only when new too. A replay the recipe refuses writes nothing. See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.
Tags
A tag is a plain string on a Run. Add one to a Run that went badly with runs.tags.update(runId, { add: ['regression-candidate'] }), or start the Run with tags. See Tag a run. --tag lists the source agent's finished Runs that carry the tag, up to 100 (runs.list with tag, like --recent, which also takes only finished Runs). A tagged Run that is still working is left out, because its first Turn has no outcome to compare a replay with yet. Salambo refuses text that could not be a tag, such as one longer than 64 characters, and the recipe says so and starts nothing. The Runs the recipe starts itself carry the tags eval and cookbook-regression-set. The calls are in src/lib/tags.ts.
What becomes a case
For each selected Run the recipe reads the first Turn from the Turn history:
const page = await salambo.runs.turns.list(runId, { limit: 1 });
const first = page.data[0];The Turn's input is the case, and the Turn's status, answer and agent_version are kept beside it as the baseline: what happened then. The Turn history is durable, so this works long after the Run's events expired.
A Run that cannot become a case is reported with its reason, never dropped without a word:
| Reason | Why |
|---|---|
| The Run has no Turns | Nothing was ever asked |
| Its first Turn had a file | A replay cannot attach the same file again |
| The Run is no longer retained | run_expired |
| No such Run in this workspace | run_not_found |
Only the first Turn is replayed. Later Turns depend on what earlier answers left in the Run's workspace, which a replay in a new Run does not have.
The replay is a new Run
Each case is replayed as a new Run on the exact agent and version you give, never by continuing the original Run. A continued Run would inherit its workspace and history, and its next Turn could follow the agent's workspace-upgrade policy onto another version.
To replay on an agent that is not production, so a replayed request never repeats a real side effect, see production and evaluation agents. The recipe reads each Run's agent with runs.retrieve and refuses to replay on that one.
Read the output
case then now movement answer now
------------- ------------ ------------ --------------- ----------------------------------------
from-run_01J v2 completed v3 completed still completes Your order ships on Friday.
from-run_01K v2 failed v3 completed now completes Refund issued. Your order ships on Frid…movement | Meaning |
|---|---|
still completes | Completed then and now |
now completes | Failed or stopped then, completes now |
regressed | Completed then, does not now. The exit code is 1 |
still does not complete | Did not complete then or now |
A replay says whether the Turn completes. Whether the answer is now right is for you, or for a check you add to the file.
The file
--out writes the cases in the format --cases reads, so the same file feeds every other recipe:
{
"cases": [
{
"id": "from-run_01K...",
"input": "Please refund order 1002.",
"baseline": {
"runId": "run_01K...",
"agentId": "agt_01P...",
"turnId": "turn_01K...",
"agentVersion": 2,
"status": "failed",
"answer": ""
}
}
]
}The file has no expected answers, because a Run only says what happened. Add mustInclude or a rubric to the cases you understand, and they become checks in task completion, reliability, compare versions and the LLM judge. Cases replayed from production can hold customer text: keep the file where you keep other customer data.