# Production and evaluation agents

Test an agent that refunds, sends or writes by deploying a second agent from the same source, wired to fake services, and checking what the fakes recorded.

An agent that refunds orders, sends email or writes to a database cannot be tested against the real thing. A test that goes wrong would refund a real customer. The pattern that solves it is two agents built from the same source.

|                          | Production agent         | Evaluation agent        |
| ------------------------ | ------------------------ | ----------------------- |
| Source                   | A commit of your project | The same commit         |
| Services                 | The real ones            | Fakes with made-up data |
| Handles                  | Real work                | Only evaluations        |
| Started by an evaluation | Never                    | Always                  |

They are two agents, not two versions of one. Each has its own ID, its own versions and its own configuration, so nothing an evaluation does can reach the production agent's configuration, credentials or Runs. Version numbers are per agent: version 3 of the evaluation agent is not version 3 of production. Compare by commit, and record which commit each version was deployed from.

## Set the evaluation agent up

Deploy the same source under a second slug, with the services pointed at fakes:

```yaml
agent:
  name: Support agent (evaluation)
  slug: support-agent-eval

env:
  ORDERS_API_URL:
    value: https://fake-orders.example.test
    description: The fake orders service. Never the real one.
    exposeTo:
      - sandbox

secrets:
  ORDERS_API_TOKEN:
    fromEnv: FAKE_ORDERS_AGENT_TOKEN
    description: The fake orders service's agent token.
    allowedHosts:
      - fake-orders.example.test
    exposeTo:
      - sandbox

runtimeConfig:
  egressPolicyMode: restricted
  egressAllowlist:
    - fake-orders.example.test
```

`salambo deploy` finds an agent by its slug and creates one if none has it, so a new slug is a new agent. The agent's tools read the address and the token from the environment instead of hard-coding them, and present the token as they would a real API key:

```js
export default function orderTools(pi) {
  pi.registerTool({
    name: 'lookup_order',
    description: 'Look up an order by its order number.',
    parameters: {
      type: 'object',
      properties: { orderId: { type: 'string' } },
      required: ['orderId'],
      additionalProperties: false,
    },
    async execute(_toolCallId, { orderId }) {
      const response = await fetch(
        `${process.env.ORDERS_API_URL}/orders/${orderId}`,
        {
          headers: { authorization: `Bearer ${process.env.ORDERS_API_TOKEN}` },
        },
      );
      return { content: [{ type: 'text', text: await response.text() }] };
    },
  });
}
```

`refund_order` is the same with `POST` and `/refund`. The agent needs a tool that reads an order by its ID: the recipe uses one to prove the agent reaches the fake. See [Environment variables and secrets](/docs/deploy/environment-secrets), [Networking and regions](/docs/deploy/networking-regions) and [Add a custom tool](/docs/agent-development/custom-tools).

The fake has to be reachable from where the evaluation agent runs. A service on your laptop is not reachable from a hosted sandbox unless you expose it, for example through a tunnel. The recipe does not deploy or configure any agent for you.

## Run it

Start the fake in one terminal:

```bash
cd examples/cookbooks
pnpm fake-orders --port 8787
```

It prints two tokens. The **agent token** is the evaluation agent's `ORDERS_API_TOKEN`. The **control token** is for the recipe, and stays out of the agent: it reads the record of calls and resets the data. Run the recipe in another terminal:

```bash
export FAKE_ORDERS_CONTROL_TOKEN="..."   # the control token it printed
pnpm prod-vs-eval \
  --agent agt_01E... --version 3 \
  --production-agent agt_01P... \
  --fake-service-url http://127.0.0.1:8787
```

`--agent` is the **evaluation** agent. `--production-agent` is only there so the recipe can refuse to evaluate against it: the recipe never starts it. See [Evaluation cookbooks](/docs/cookbooks/overview) for the key and the address of your stack.

## Checks before any case that writes

The recipe stops with an explanation, before it starts a Run that could write, if any of these fails:

1. **The evaluation agent is not the production agent.** Same ID, and it refuses.
2. **The address is a fake, and says so.** The recipe asks `GET /health` and needs `{"service": "fake-orders", "mode": "fake"}`. Anything else at that address, including the real orders service, is refused, and nothing but the health check is sent to it. The recipe reads and resets whatever is at the address, so it must only ever be the fake.
3. **The fake is running.** If nothing answers, the message says how to start it.
4. **The agent reaches the fake.** The health check proves the address is a fake. It cannot prove where the evaluation agent sends its requests: an agent still pointed at the real orders service would pass it. So the recipe starts one read-only Run first, asking the agent to look up order `FAKE-CANARY`, an order that exists only in the fake. If the fake did not record that lookup, the recipe stops with the Run's ID and starts nothing that writes. An agent pointed at the real service sent its lookup there, where a made-up order number finds nothing.

The canary proves the agent reaches the fake for reads. It cannot prove that every service the agent uses is a fake. Wire every service the agent can write to.

## What the fake is for

The fake orders service has the routes of the real one, made-up orders, and a record of every call it receives:

| Route                     | Token   | Does                                                           |
| ------------------------- | ------- | -------------------------------------------------------------- |
| `GET /orders/:id`         | Agent   | Returns an order, or 404                                       |
| `POST /orders/:id/refund` | Agent   | Refunds a paid order, and answers 409 for one already refunded |
| `GET /health`             | None    | Says that this is the fake                                     |
| `GET /__calls`            | Control | Every call the agent made                                      |
| `POST /__reset`           | Control | Restores the made-up data and clears the record                |

Its orders are `1001` (paid), `1002` (paid), `1003` (already refunded) and `FAKE-CANARY`. The recipe resets it before each case, so every case starts from the same data.

The record is the point, so it has to be trustworthy when the fake can be reached by more than the agent. A call to an order route without the agent token is refused and is not recorded, and the control routes need the control token, so nobody else can add to the record, read it or wipe it. The tokens are random unless you pass `--agent-token` and `--control-token`, and the fake never reads them from the environment, where a real service's token could be mistaken for a fake's. Bind it to `127.0.0.1`, its default, unless the agent must reach it from elsewhere.

A test that rests on what the agent says it did is a test of its wording. The recipe asks the fake which writes really happened, and compares them with the writes each case expects:

```ts
export const ACTING_CASES: readonly ActingCase[] = [
  {
    id: 'refunds-a-paid-order',
    input:
      'Grace Hopper asks for a refund of order 1002. Process it and tell her it is done.',
    mustInclude: ['1002'],
    writes: ['POST /orders/1002/refund'],
  },
  {
    id: 'does-not-refund-twice',
    input: 'Please refund order 1003.',
    mustInclude: ['already'],
    writes: [],
  },
];
```

`writes` is every call that changes something and is expected to have happened, each once. Reads are not listed: an agent may look at as many orders as it needs. A refund the agent claims but the fake never saw is **missing**. A refund the fake saw that the case did not ask for, or a second one for the same order, is **unexpected**.

## Read the output

```text
case                    turn       answer  writes on the fake  detail
----------------------  ---------  ------  ------------------  -----------------------------------
refunds-a-paid-order    completed  ok      not as expected     missing POST /orders/1002/refund
does-not-refund-twice   completed  ok      not as expected     unexpected POST /orders/1003/refund
reads-without-changing  completed  ok      as expected         none
```

In the first row the agent said it refunded order 1002 and the fake never saw it. In the second it refunded order 1003 without looking, and the fake refused. The exit code is 1 when any answer or any write was not as expected.

The cases run one at a time, because they share one fake and reset it. Give each case its own fake to run them side by side.

## Keep the fake honest

A fake is only as true as it is kept. When the real service changes a route, a status code or a rule, change the fake in the same commit. The cases in this recipe are tied to the fake's made-up data, so they live in `src/09-prod-vs-eval.ts`, next to it. Copy the fake for your own services: it is one short file, and the pattern is the same for a mail service, a database or a payment provider.
