# Testing and publishing

Try the agent before customers do, score it against cases you wrote, and put it live deliberately.

Source: https://docs.omazy.ai/how-to/agent/testing/

Three console surfaces sit between a change and a customer: **Preview**,
**Testing and Evals**, and **Publish**. They do different jobs and it is worth
knowing which one answers which question.

| Surface | Answers |
|---|---|
| **Preview** | What does this feel like to talk to? |
| **Testing and Evals** | Does it still get the answers right? |
| **Publish** | Is it live, and what changed? |

## Preview

A conversation with your unpublished agent. Use it the way a customer would:
type badly, change your mind halfway, ask the awkward pricing question.

Preview is for judgement, not proof. It tells you whether the tone is right and
whether the agent rambles. It cannot tell you that yesterday's fix did not break
last week's answer, because you will not remember to check.

## Testing and Evals

This is where the checking becomes repeatable. You write **cases**, and a case
has three parts.

**User turns.** What the customer says. One line for a simple check, several for
a conversation where the agent has to hold context.

**Assertions.** What has to be true about the reply. This is the part that turns
a vibe into a test. "Mentions the deposit amount" is an assertion. "Sounds good"
is not.

**Seed context.** Memory the agent should already have, so you can test how it
behaves for a returning customer without staging a real one.

Run the suite and each case is scored, with the reply and the reason kept so you
can open a case and read what actually happened rather than guessing from a
number.

### Write cases from real failures

The temptation is to write cases for the things you know work. Those pass
forever and teach you nothing.

The cases worth having come from the inbox. Every handover where the agent
should have coped is a case waiting to be written, already phrased the way a real
customer phrased it. Turn the fix into a case at the same time you fix it, and
that bug cannot come back quietly.

### Comparing models

A case can be run across models, and you can open one to read what each replied.

This is the honest way to answer "should we switch models", which is otherwise
decided by whoever read a benchmark most recently. The bigger model is not
automatically better at following your brief, and in live chat a slower correct
answer loses to a fast good one more often than people admit.

## Publish

Publishing takes the agent you have been editing and makes it the one answering
customers. Until you press it, your edits are not live. That is what lets you
leave a half-finished brief overnight.

Publish also keeps release history, so you can see what went live and when. Every
publish of the brief stores a version, which means rolling back is restoring a
version rather than reconstructing one from memory.

An agent can be unpublished as well as published. Unpublishing is the blunt
instrument for "stop answering right now", and it is the correct move when
something is badly wrong. Fixing a live agent while it is still talking to people
is how a small problem becomes a long afternoon.

## The order that works

1. Change one thing.
2. Preview it, to see whether it reads right.
3. Run the suite, to see whether anything else broke.
4. Publish.
5. Add a case for whatever you just fixed.

Step 5 is the one everyone skips and the only one that compounds.
