Salambo
Browse documentation
Evaluation cookbooksRegression set from tagged Runs

Regression set from tagged Runs

Collect Runs that went wrong, read what each was asked from its Turn history, write them as cases, and replay them on an exact version.

View Markdown

Production is where the cases you did not think of turn up. When someone finds a Run that went badly, they tag it with a plain string such as regression-candidate. This recipe collects those Runs, turns each one into a case, writes the cases to a file, and replays them on a version you name. You learn whether a fix fixed it, and whether anything that used to work stopped.

Run it

By Run ID, which works today:

bash
cd examples/cookbooks
pnpm regression-set \
  --run-ids run_01J...,run_01K... \
  --agent agt_01L... --version 3 \
  --out regression-set.json

By the most recent finished Runs of an agent, when you have no IDs yet:

bash
pnpm regression-set --recent 20 --source-agent agt_01J... \
  --agent agt_01L... --version 3

By tag:

bash
pnpm regression-set --tag regression-candidate --source-agent agt_01J... \
  --agent agt_01L... --version 3

Replay goes to a separate evaluation agent: the recipe refuses to replay a Run's request on the agent that Run came from, because that repeats whatever the agent did the first time, such as a refund or an email. Pass --allow-source-agent only for an agent with no side effects.

Choose the Runs with exactly one of --tag, --run-ids and --recent. --recent takes up to 100 finished Runs of the source agent, including evaluation Runs, so point it at an agent that only serves real traffic. With --no-replay the recipe only writes the file. It never replaces a file that is already there, because that file may hold checks you added by hand: choose another --out, or pass --overwrite, which writes the whole replacement beside the file before moving it into place, so a failed write leaves the old file as it was, and never leaves the replacement more open than the file was. A file that others could read is replaced by one that only its owner can read, and the recipe says so. The file holds what Runs were asked, which can be customer text, so a new one is readable by its owner alone. A symbolic link is refused: pass the path of the file itself. The --json results file is owner-only when new too. A replay the recipe refuses writes nothing. See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.

Tags

A tag is a plain string on a Run. Add one to a Run that went badly with runs.tags.update(runId, { add: ['regression-candidate'] }), or start the Run with tags. See Tag a run. --tag lists the source agent's finished Runs that carry the tag, up to 100 (runs.list with tag, like --recent, which also takes only finished Runs). A tagged Run that is still working is left out, because its first Turn has no outcome to compare a replay with yet. Salambo refuses text that could not be a tag, such as one longer than 64 characters, and the recipe says so and starts nothing. The Runs the recipe starts itself carry the tags eval and cookbook-regression-set. The calls are in src/lib/tags.ts.

What becomes a case

For each selected Run the recipe reads the first Turn from the Turn history:

ts
const page = await salambo.runs.turns.list(runId, { limit: 1 });
const first = page.data[0];

The Turn's input is the case, and the Turn's status, answer and agent_version are kept beside it as the baseline: what happened then. The Turn history is durable, so this works long after the Run's events expired.

A Run that cannot become a case is reported with its reason, never dropped without a word:

ReasonWhy
The Run has no TurnsNothing was ever asked
Its first Turn had a fileA replay cannot attach the same file again
The Run is no longer retainedrun_expired
No such Run in this workspacerun_not_found

Only the first Turn is replayed. Later Turns depend on what earlier answers left in the Run's workspace, which a replay in a new Run does not have.

The replay is a new Run

Each case is replayed as a new Run on the exact agent and version you give, never by continuing the original Run. A continued Run would inherit its workspace and history, and its next Turn could follow the agent's workspace-upgrade policy onto another version.

To replay on an agent that is not production, so a replayed request never repeats a real side effect, see production and evaluation agents. The recipe reads each Run's agent with runs.retrieve and refuses to replay on that one.

Read the output

text
case           then          now           movement         answer now
-------------  ------------  ------------  ---------------  ----------------------------------------
from-run_01J   v2 completed  v3 completed  still completes  Your order ships on Friday.
from-run_01K   v2 failed     v3 completed  now completes    Refund issued. Your order ships on Frid…
movementMeaning
still completesCompleted then and now
now completesFailed or stopped then, completes now
regressedCompleted then, does not now. The exit code is 1
still does not completeDid not complete then or now

A replay says whether the Turn completes. Whether the answer is now right is for you, or for a check you add to the file.

The file

--out writes the cases in the format --cases reads, so the same file feeds every other recipe:

json
{
  "cases": [
    {
      "id": "from-run_01K...",
      "input": "Please refund order 1002.",
      "baseline": {
        "runId": "run_01K...",
        "agentId": "agt_01P...",
        "turnId": "turn_01K...",
        "agentVersion": 2,
        "status": "failed",
        "answer": ""
      }
    }
  ]
}

The file has no expected answers, because a Run only says what happened. Add mustInclude or a rubric to the cases you understand, and they become checks in task completion, reliability, compare versions and the LLM judge. Cases replayed from production can hold customer text: keep the file where you keep other customer data.