Browse documentation
Production and evaluation agents
Test an agent that refunds, sends or writes by deploying a second agent from the same source, wired to fake services, and checking what the fakes recorded.
View MarkdownAn agent that refunds orders, sends email or writes to a database cannot be tested against the real thing. A test that goes wrong would refund a real customer. The pattern that solves it is two agents built from the same source.
| Production agent | Evaluation agent | |
|---|---|---|
| Source | A commit of your project | The same commit |
| Services | The real ones | Fakes with made-up data |
| Handles | Real work | Only evaluations |
| Started by an evaluation | Never | Always |
They are two agents, not two versions of one. Each has its own ID, its own versions and its own configuration, so nothing an evaluation does can reach the production agent's configuration, credentials or Runs. Version numbers are per agent: version 3 of the evaluation agent is not version 3 of production. Compare by commit, and record which commit each version was deployed from.
Set the evaluation agent up
Deploy the same source under a second slug, with the services pointed at fakes:
agent:
name: Support agent (evaluation)
slug: support-agent-eval
env:
ORDERS_API_URL:
value: https://fake-orders.example.test
description: The fake orders service. Never the real one.
exposeTo:
- sandbox
secrets:
ORDERS_API_TOKEN:
fromEnv: FAKE_ORDERS_AGENT_TOKEN
description: The fake orders service's agent token.
allowedHosts:
- fake-orders.example.test
exposeTo:
- sandbox
runtimeConfig:
egressPolicyMode: restricted
egressAllowlist:
- fake-orders.example.testsalambo deploy finds an agent by its slug and creates one if none has it, so a new slug is a new agent. The agent's tools read the address and the token from the environment instead of hard-coding them, and present the token as they would a real API key:
export default function orderTools(pi) {
pi.registerTool({
name: 'lookup_order',
description: 'Look up an order by its order number.',
parameters: {
type: 'object',
properties: { orderId: { type: 'string' } },
required: ['orderId'],
additionalProperties: false,
},
async execute(_toolCallId, { orderId }) {
const response = await fetch(
`${process.env.ORDERS_API_URL}/orders/${orderId}`,
{
headers: { authorization: `Bearer ${process.env.ORDERS_API_TOKEN}` },
},
);
return { content: [{ type: 'text', text: await response.text() }] };
},
});
}refund_order is the same with POST and /refund. The agent needs a tool that reads an order by its ID: the recipe uses one to prove the agent reaches the fake. See Environment variables and secrets, Networking and regions and Add a custom tool.
The fake has to be reachable from where the evaluation agent runs. A service on your laptop is not reachable from a hosted sandbox unless you expose it, for example through a tunnel. The recipe does not deploy or configure any agent for you.
Run it
Start the fake in one terminal:
cd examples/cookbooks
pnpm fake-orders --port 8787It prints two tokens. The agent token is the evaluation agent's ORDERS_API_TOKEN. The control token is for the recipe, and stays out of the agent: it reads the record of calls and resets the data. Run the recipe in another terminal:
export FAKE_ORDERS_CONTROL_TOKEN="..." # the control token it printed
pnpm prod-vs-eval \
--agent agt_01E... --version 3 \
--production-agent agt_01P... \
--fake-service-url http://127.0.0.1:8787--agent is the evaluation agent. --production-agent is only there so the recipe can refuse to evaluate against it: the recipe never starts it. See Evaluation cookbooks for the key and the address of your stack.
Checks before any case that writes
The recipe stops with an explanation, before it starts a Run that could write, if any of these fails:
- The evaluation agent is not the production agent. Same ID, and it refuses.
- The address is a fake, and says so. The recipe asks
GET /healthand needs{"service": "fake-orders", "mode": "fake"}. Anything else at that address, including the real orders service, is refused, and nothing but the health check is sent to it. The recipe reads and resets whatever is at the address, so it must only ever be the fake. - The fake is running. If nothing answers, the message says how to start it.
- The agent reaches the fake. The health check proves the address is a fake. It cannot prove where the evaluation agent sends its requests: an agent still pointed at the real orders service would pass it. So the recipe starts one read-only Run first, asking the agent to look up order
FAKE-CANARY, an order that exists only in the fake. If the fake did not record that lookup, the recipe stops with the Run's ID and starts nothing that writes. An agent pointed at the real service sent its lookup there, where a made-up order number finds nothing.
The canary proves the agent reaches the fake for reads. It cannot prove that every service the agent uses is a fake. Wire every service the agent can write to.
What the fake is for
The fake orders service has the routes of the real one, made-up orders, and a record of every call it receives:
| Route | Token | Does |
|---|---|---|
GET /orders/:id | Agent | Returns an order, or 404 |
POST /orders/:id/refund | Agent | Refunds a paid order, and answers 409 for one already refunded |
GET /health | None | Says that this is the fake |
GET /__calls | Control | Every call the agent made |
POST /__reset | Control | Restores the made-up data and clears the record |
Its orders are 1001 (paid), 1002 (paid), 1003 (already refunded) and FAKE-CANARY. The recipe resets it before each case, so every case starts from the same data.
The record is the point, so it has to be trustworthy when the fake can be reached by more than the agent. A call to an order route without the agent token is refused and is not recorded, and the control routes need the control token, so nobody else can add to the record, read it or wipe it. The tokens are random unless you pass --agent-token and --control-token, and the fake never reads them from the environment, where a real service's token could be mistaken for a fake's. Bind it to 127.0.0.1, its default, unless the agent must reach it from elsewhere.
A test that rests on what the agent says it did is a test of its wording. The recipe asks the fake which writes really happened, and compares them with the writes each case expects:
export const ACTING_CASES: readonly ActingCase[] = [
{
id: 'refunds-a-paid-order',
input:
'Grace Hopper asks for a refund of order 1002. Process it and tell her it is done.',
mustInclude: ['1002'],
writes: ['POST /orders/1002/refund'],
},
{
id: 'does-not-refund-twice',
input: 'Please refund order 1003.',
mustInclude: ['already'],
writes: [],
},
];writes is every call that changes something and is expected to have happened, each once. Reads are not listed: an agent may look at as many orders as it needs. A refund the agent claims but the fake never saw is missing. A refund the fake saw that the case did not ask for, or a second one for the same order, is unexpected.
Read the output
case turn answer writes on the fake detail
---------------------- --------- ------ ------------------ -----------------------------------
refunds-a-paid-order completed ok not as expected missing POST /orders/1002/refund
does-not-refund-twice completed ok not as expected unexpected POST /orders/1003/refund
reads-without-changing completed ok as expected noneIn the first row the agent said it refunded order 1002 and the fake never saw it. In the second it refunded order 1003 without looking, and the fake refused. The exit code is 1 when any answer or any write was not as expected.
The cases run one at a time, because they share one fake and reset it. Give each case its own fake to run them side by side.
Keep the fake honest
A fake is only as true as it is kept. When the real service changes a route, a status code or a rule, change the fake in the same commit. The cases in this recipe are tied to the fake's made-up data, so they live in src/09-prod-vs-eval.ts, next to it. Copy the fake for your own services: it is one short file, and the pattern is the same for a mail service, a database or a payment provider.