Salambo
Browse documentation
Evaluation cookbooksCost and latency against completion

Cost and latency against completion

Put completion, latency and tokens side by side for each exact version, without folding them into one score.

View Markdown

A version that answers correctly more often but takes three times as long, or spends twice the tokens, is a trade, not a win. This recipe measures completion, latency and tokens for each version you name and prints them side by side. It stops there: it does not rank the versions, because the exchange rate between a second, a token and a wrong answer is yours. A nightly batch cares about tokens. A chat cares about seconds.

Run it

bash
cd examples/cookbooks
pnpm cost-latency --agent agt_01J... --versions 3,4 --attempts 3 --price-per-million-tokens 3

See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.

FlagMeaning
--versionsThe exact versions to measure, for example 3,4
--attemptsAttempts per case on each version. Default 3
--price-per-million-tokensOptional. Adds an estimate of what the median Run's tokens cost at your price
--completion-onlyAccept cases with no check and measure only whether their Turns complete

Where each figure comes from

FigureSourceDurable?
completed, passedThe Turn's status, and the case's check on its answerYes
latency p50, latency p95duration_ms of the Turns that finished, as the platform measured themYes
tokens p50The usage the model reported on its finished messages, summed per RunNo: events
tool calls p50The tools the agent startedNo: events

Tokens and tool calls come from a Run's events, which are a rolling window. The recipe only counts them when the events are complete. When a Run's events are partial, expired, empty or could not be read, its tokens are missing, not zero, and the table says how many Runs the figure rests on: 4700 (9/9 Runs). A median over the Runs that had them is honest. A median with zeros in for the ones that did not is not.

text
version  completed  passed  latency p50  latency p95  tokens p50        tool calls p50  est. per Run
-------  ---------  ------  -----------  -----------  ----------------  --------------  ------------
v2       9/9        3/9     4.1s         4.1s         4700 (9/9 Runs)   0               0.0141
v3       9/9        9/9     12.4s        12.4s        12300 (9/9 Runs)  2               0.0369

Read the columns against each other: here v3 is right three times as often and takes three times as long, and whether that is worth it depends on what the agent is for.

Tokens are not a bill

Tokens are what the runtime reported in retained events. The public API returns no cost figures, and the estimate is your median tokens times a price you supplied, in whatever currency you use. What Salambo charges is on the workspace's usage and billing pages, which are the source of truth. See Billing and usage.

How it works

Latency is duration_ms from the Turn record. Percentiles are nearest-rank, so each is a value that was measured and never an interpolation, and latency counts only Turns that finished. Runs of the versions alternate in the queue, so a slow hour lands on all of them.

ts
const evidence = await readEvidence(
  salambo,
  observation.runId,
  observation.turnId,
);
const tokens = isConclusive(evidence) ? tokenUsage(evidence.events) : null;

tokenUsage sums usage over the assistant messages that finished (message_end events) and returns null when none carried usage, so absent is never zero.

Change it

Measure with cases that look like your traffic: the median of three toy questions says little about a task that reads a repository. Add --attempts for steadier percentiles. With few Runs, p95 is just the slowest one.