Browse documentation
Cost and latency against completion
Put completion, latency and tokens side by side for each exact version, without folding them into one score.
View MarkdownA version that answers correctly more often but takes three times as long, or spends twice the tokens, is a trade, not a win. This recipe measures completion, latency and tokens for each version you name and prints them side by side. It stops there: it does not rank the versions, because the exchange rate between a second, a token and a wrong answer is yours. A nightly batch cares about tokens. A chat cares about seconds.
Run it
cd examples/cookbooks
pnpm cost-latency --agent agt_01J... --versions 3,4 --attempts 3 --price-per-million-tokens 3See Evaluation cookbooks for the key, the address of your stack and the flags every recipe takes.
| Flag | Meaning |
|---|---|
--versions | The exact versions to measure, for example 3,4 |
--attempts | Attempts per case on each version. Default 3 |
--price-per-million-tokens | Optional. Adds an estimate of what the median Run's tokens cost at your price |
--completion-only | Accept cases with no check and measure only whether their Turns complete |
Where each figure comes from
| Figure | Source | Durable? |
|---|---|---|
completed, passed | The Turn's status, and the case's check on its answer | Yes |
latency p50, latency p95 | duration_ms of the Turns that finished, as the platform measured them | Yes |
tokens p50 | The usage the model reported on its finished messages, summed per Run | No: events |
tool calls p50 | The tools the agent started | No: events |
Tokens and tool calls come from a Run's events, which are a rolling window. The recipe only counts them when the events are complete. When a Run's events are partial, expired, empty or could not be read, its tokens are missing, not zero, and the table says how many Runs the figure rests on: 4700 (9/9 Runs). A median over the Runs that had them is honest. A median with zeros in for the ones that did not is not.
version completed passed latency p50 latency p95 tokens p50 tool calls p50 est. per Run
------- --------- ------ ----------- ----------- ---------------- -------------- ------------
v2 9/9 3/9 4.1s 4.1s 4700 (9/9 Runs) 0 0.0141
v3 9/9 9/9 12.4s 12.4s 12300 (9/9 Runs) 2 0.0369Read the columns against each other: here v3 is right three times as often and takes three times as long, and whether that is worth it depends on what the agent is for.
Tokens are not a bill
Tokens are what the runtime reported in retained events. The public API returns no cost figures, and the estimate is your median tokens times a price you supplied, in whatever currency you use. What Salambo charges is on the workspace's usage and billing pages, which are the source of truth. See Billing and usage.
How it works
Latency is duration_ms from the Turn record. Percentiles are nearest-rank, so each is a value that was measured and never an interpolation, and latency counts only Turns that finished. Runs of the versions alternate in the queue, so a slow hour lands on all of them.
const evidence = await readEvidence(
salambo,
observation.runId,
observation.turnId,
);
const tokens = isConclusive(evidence) ? tokenUsage(evidence.events) : null;tokenUsage sums usage over the assistant messages that finished (message_end events) and returns null when none carried usage, so absent is never zero.
Change it
Measure with cases that look like your traffic: the median of three toy questions says little about a task that reads a repository. Add --attempts for steadier percentiles. With few Runs, p95 is just the slowest one.