Salambo
Browse documentation
Evaluation cookbooksCompare two exact versions

Compare two exact versions

Run the same cases on two versions of one agent, side by side, and report per case which way each moved.

View Markdown

Is this version better than that one? This recipe starts the same cases on a baseline version and a candidate version of one agent and reports, case by case, whether the candidate passed more attempts, fewer, or as many.

Run it

bash
cd examples/cookbooks
pnpm compare-versions --agent agt_01J... --baseline 3 --candidate 4 --attempts 3

See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.

Every case needs a mustInclude or mustNotInclude to have an answer to compare. A case with neither is refused, unless you pass --completion-only to compare only whether Turns complete.

Both versions are named by number. Starting a Run on an exact version does not activate it, so the agent's users keep the version they have while the candidate is tested. See Start on an exact agent version.

What makes it a fair comparison

  • Same cases, same checks, same attempts for both versions.
  • Side by side in the queue. The attempts of the baseline and the candidate alternate, so a slow hour or a busy provider lands on both.
  • Every Turn is checked for the version that ran it. A Turn reports agent_version, pinned when it was admitted. If one differs from the version asked for, the recipe stops and prints nothing scored, because the numbers would be about another version.
  • Every start is a new Run. The recipe never continues a Run with run_id. A later Turn of an existing Run follows the agent's workspace-upgrade policy and can move to a different version, which would mix the two.
ts
const observed = await mapWithLimit(
  planRuns(cases, [baseline, candidate], attempts),
  limits.concurrency,
  ({ evalCase, target, attempt }) =>
    observer.observe(target, evalCase, attempt),
);

planRuns (in src/lib/attempts.ts) lists the targets of one attempt next to each other, which is what puts the two versions side by side in the queue.

Read the output

text
case        v3 passed  v4 passed  movement   errored (both)
----------  ---------  ---------  ---------  --------------
arithmetic  3/3        0/3        regressed  0, 0
json-shape  0/3        3/3        improved   0, 0
summary     3/3        3/3        unchanged  0, 0

movement compares the passes of the two versions on that case: improved, regressed or unchanged. The last column counts Turns that did not complete on each side, since a version that fails to complete is a different problem from one that answers wrongly. The exit code is 1 when the candidate regressed on any case.

No total

The recipe does not add the cases up. A candidate that gains on ten cases and loses on one is not "+9": the one it lost may be the case your customers hit most. Read the regressions first.

Is the difference real?

With three attempts, a difference of one attempt can be chance. The recipe says so, and the honest next step is to run the cases that moved through the reliability recipe on each version with more attempts. Compare versions to find where to look. Use more attempts to decide.

Change it

The versions are flags, so a comparison is a command. To compare an agent against a later version of itself over your real cases, pass --cases. To compare cost and speed as well as pass rate, use cost and latency. If the two versions belong to different agents, run the recipe once per agent and compare the outputs yourself: version numbers are per agent.