Pick the right tool for the job
The pattern is the one Inbox is built around: find a failure, trace it in AI Steps, fix it at the source, replay to confirm, then lock it in as a Test Suite test case.
How Test Suite works
Every test case in Test Suite starts from a real conversation. A run replays the customer’s messages from that conversation, generates new agent replies, and evaluates them against the case’s criteria. Test Suite does not invent new customer scenarios: to test another customer path, create or select another conversation.Test Suite uses recorded or mock responses for external Actions. There is no live Action execution in Test Suite, and running with draft Behavior, rules, or articles does not publish those drafts. For live end-to-end checks of an Action, use a sandbox environment or a controlled pilot.
What to test
Cover each of these categories for every agent. Most teams start with the first three and add the rest within the first weeks.
Use an AI judgement when the outcome needs interpretation, such as whether a reply explains a refund policy without inventing an exception. Use an exact check when the outcome is an observable event, such as whether the agent called a particular Action or handed off to a human. Exact checks don’t vary between runs, so lean on them for anything that must never break.
Organize cases, groups, and collections
Keep each criteria group focused on one behavior. A failing run then tells you immediately what regressed. A small, well-chosen set of cases beats a large random one.
A criteria group’s conversation picker can add up to 50 conversations at a time from Inbox. For a single conversation, use the flask icon in Inbox, select it, and click Select to open the case editor.
Then group by release scope: put the cases and criteria groups for one feature or launch into a collection, and use Run collection before that feature changes. Use All tests for the full regression pass.
Write each AI judgement as the outcome you want across the whole conversation, not exact wording. For example: “Agent verifies the account, explains the refund window, and escalates to billing only if the order is outside it.”
Editing a criteria group’s criteria changes what its linked cases use on their next run. Saved run results do not change. Review the affected cases before confirming a group update.
Run the regression check before every change
Make this the release gate for any change to a prompt, knowledge source, Rulebook rule, Action, or model.1
Test the change in isolation
For a rule change, run Suggested scenarios on the draft. For a knowledge or prompt change, open the conversations it should affect in Inbox and replay them against the version you’re testing.
2
Run the affected cases against your drafts
You don’t need to publish first. In the run dialog, under Behaviour, choose Latest published behaviour or a saved version, and select the rule and article drafts you want to test. Selected drafts replace their published versions for that run only. Nothing is published. Use Run collection for the affected scope or Run all for everything.
3
Read the verdicts, not just the count
Open Runs. Separate Failed (a required criterion did not pass) from Execution error (the run could not complete normally). Expand each failed result and read the criterion-level reasoning, the generated conversation, and the Action responses.
4
Fix, then re-run
Fix at the source, then use Rerun test case or run the collection again. If an execution error reports a missing Action response, click Add mock response, save it, and rerun.
5
Publish the change
Once the required criteria pass, publish the prompt, rule, or article drafts you tested.
6
Watch production after release
Check Analytics the following days for changes in AI resolution rate, escalation reasons, and CSAT. See Measure your resolution rate.
Habits that keep tests useful
- Add every fixed bug as a test case. Before you close the loop on a wrong reply, use the flask icon in Inbox to create a case from that conversation and write the criterion that would have caught it. Your suite then grows with every failure you learn about.
- Change one variable at a time. If you edit a prompt and a rule together, a changed verdict can’t tell you which one caused it.
- Check the run configuration before blaming the agent. Compare the recorded criteria, test setup, and selected draft versions with the earlier run. A changed verdict can come from an edited criterion or setup rather than a different agent.
- Write AI judgements as outcomes. Expectations that describe the outcome you want are stable across runs; expectations that demand exact wording fail on harmless rephrasing.
- Separate errors from failures. An execution error means the case could not complete normally, often because a recorded Action response is missing. Fix the setup before you read the verdict.
- Test both Action outcomes. Use Duplicate test case to copy a case, then set one copy’s mock response to Success and the other to Failure. The copy’s setup and case-specific criteria are independent; linked group criteria stay shared.
- Keep required checks few and sharp. Mark only the criteria that define the verdict as Required to pass. Non-required checks still show in results without failing the case.
Automate parts of it with the API
Test Suite runs are started from the dashboard. The public API for the rebuilt Test Suite is not documented yet, so build CI checks around the endpoints below.
Replays create separate replay conversations. They don’t overwrite the original conversation or apply any changes by themselves. See Replays for modes and model overrides, and API Keys for scopes.
Related
Test Suite
Test cases, criteria groups, collections, and runs.
Inbox
AI Steps, replays, feedback, and creating test cases from conversations.
Refine with AI
Turn a wrong reply into a reviewed fix.
Intent Rules
Suggested scenarios and manual test runs for each rule.

