# Cost and latency against completion

Put completion, latency and tokens side by side for each exact version, without folding them into one score.

A version that answers correctly more often but takes three times as long, or spends twice the tokens, is a trade, not a win. This recipe measures completion, latency and tokens for each version you name and prints them side by side. It stops there: it does not rank the versions, because the exchange rate between a second, a token and a wrong answer is yours. A nightly batch cares about tokens. A chat cares about seconds.

## Run it

```bash
cd examples/cookbooks
pnpm cost-latency --agent agt_01J... --versions 3,4 --attempts 3 --price-per-million-tokens 3
```

See [Evaluation cookbooks](/docs/cookbooks/overview) for the key, the address of your stack and the flags every recipe takes.

| Flag                         | Meaning                                                                       |
| ---------------------------- | ----------------------------------------------------------------------------- |
| `--versions`                 | The exact versions to measure, for example `3,4`                              |
| `--attempts`                 | Attempts per case on each version. Default 3                                  |
| `--price-per-million-tokens` | Optional. Adds an estimate of what the median Run's tokens cost at your price |
| `--completion-only`          | Accept cases with no check and measure only whether their Turns complete      |

## Where each figure comes from

| Figure                       | Source                                                                  | Durable?   |
| ---------------------------- | ----------------------------------------------------------------------- | ---------- |
| `completed`, `passed`        | The Turn's status, and the case's check on its answer                   | Yes        |
| `latency p50`, `latency p95` | `duration_ms` of the Turns that finished, as the platform measured them | Yes        |
| `tokens p50`                 | The usage the model reported on its finished messages, summed per Run   | No: events |
| `tool calls p50`             | The tools the agent started                                             | No: events |

Tokens and tool calls come from a Run's events, which are a rolling window. The recipe only counts them when the events are complete. When a Run's events are `partial`, `expired`, `empty` or could not be read, its tokens are **missing, not zero**, and the table says how many Runs the figure rests on: `4700 (9/9 Runs)`. A median over the Runs that had them is honest. A median with zeros in for the ones that did not is not.

```text
version  completed  passed  latency p50  latency p95  tokens p50        tool calls p50  est. per Run
-------  ---------  ------  -----------  -----------  ----------------  --------------  ------------
v2       9/9        3/9     4.1s         4.1s         4700 (9/9 Runs)   0               0.0141
v3       9/9        9/9     12.4s        12.4s        12300 (9/9 Runs)  2               0.0369
```

Read the columns against each other: here `v3` is right three times as often and takes three times as long, and whether that is worth it depends on what the agent is for.

## Tokens are not a bill

Tokens are what the runtime reported in retained events. The public API returns no cost figures, and the estimate is your median tokens times a price you supplied, in whatever currency you use. What Salambo charges is on the workspace's usage and billing pages, which are the source of truth. See [Billing and usage](/docs/concepts/billing-usage).

## How it works

Latency is `duration_ms` from the Turn record. Percentiles are nearest-rank, so each is a value that was measured and never an interpolation, and latency counts only Turns that finished. Runs of the versions alternate in the queue, so a slow hour lands on all of them.

```ts
const evidence = await readEvidence(
  salambo,
  observation.runId,
  observation.turnId,
);
const tokens = isConclusive(evidence) ? tokenUsage(evidence.events) : null;
```

`tokenUsage` sums `usage` over the assistant messages that finished (`message_end` events) and returns `null` when none carried usage, so absent is never zero.

## Change it

Measure with cases that look like your traffic: the median of three toy questions says little about a task that reads a repository. Add `--attempts` for steadier percentiles. With few Runs, p95 is just the slowest one.
