Start from real tickets, not invented questions
Questions your team writes from memory are cleaner than what customers actually send. Pull a sample of recent closed tickets from your helpdesk and build the test cases from them, keeping the customer’s original wording (with personal data removed). Choose cases along three axes so they reflect your real workload and your real exposure:
A low-volume, high-risk intent (for example, a complaint that may need formal handling) deserves cases even if it appears rarely. Volume tells you where most conversations land; risk tells you where a single failure hurts.
Write golden answers as outcomes, not scripts
A golden answer is the reference for what a correct response looks like. Write it as the outcome you expect across the conversation, not the exact words. AI agents phrase things differently each run, so word-for-word expectations fail for reasons that don’t matter. A useful golden answer has four parts:- The facts that must appear. The refund window, the fee amount, the eligibility condition.
- The source those facts come from. The specific approved article, policy or system field.
- The action, if any. Which operation should run, with which inputs.
- The end state. Resolved, waiting for the customer, or handed to a human with context.
If you use an automated grader (an AI model that evaluates replies against written criteria), calibrate it against human grading on a sample first, and keep a human review of a share of results during the pilot period.
Cover the edge cases that break agents
Happy paths prove the agent works. Edge cases show you where it fails. Include each of these:Catch wrong policy answers
The most damaging failure is a confident, fluent answer that states the wrong policy. It reads fine, so casual review misses it. Test for it on purpose.- Require traceability. For every policy case, the grader checks that each claim comes from an approved source. An answer that is correct by luck but not grounded is a fail, because it will be wrong next time.
- Test conflicting-policy cases. If two documents disagree (an old help article and a newer internal policy, or two regions with different rules), write cases that hit the conflict. The expected outcome is that the agent uses the source you designated as authoritative, or escalates. Fix the conflict in the knowledge itself afterwards.
- Test stale content. Include a case for any policy that changed recently. If the agent quotes the old version, find and retire the old source.
- Test the gap. Ask something your knowledge doesn’t cover. The right answer is “I don’t have that information” plus a handoff, not a plausible guess.
Test actions in a sandbox
If the agent can take actions (issue refunds, change plans, update addresses), test them against a sandbox or staging environment, never production. Check that:- The right action runs only when its conditions are met, and doesn’t run when they aren’t.
- Inputs are correct: the right account, amount and item.
- Failures are handled. Make the sandbox return an error or a timeout and confirm the agent tells the customer accurately and hands off.
- Confirmation is asked for where your policy requires it, before anything irreversible.
Test escalation both ways
Write cases that must escalate (explicit request for a human, legal threat, vulnerable customer, compliance trigger) and cases that must not (a customer venting about a delay the agent can resolve). For each handoff, check that the human receives enough context to continue without asking the customer to repeat themselves.Run the loop on every change
Testing is not a one-time gate. Every change to a prompt, knowledge source, workflow, action or model can break something that used to work. Rerun the cases each time and compare against previous runs. Two habits keep the loop useful. First, change one thing at a time, so a drop in results points to one cause. Second, add every real failure you find, in testing or in production, as a new case. Your cases then grow with what you learn.Set pass thresholds as a team decision
There is no universal score that means “ready”. Agree on thresholds before the first run, with support, compliance and product in the room, and write them down. A common pattern is to set them per category rather than overall:- Zero tolerance categories, such as wrong policy statements on regulated topics, unauthorized actions or data shown to the wrong person. Any failure blocks launch.
- High bar categories, such as top-volume intents and escalation cases.
- Improving categories, such as tone or rare intents, where you accept some misses and track the trend.
Sample test-case table
Use a simple table (a spreadsheet works) as the source of truth for your cases. The rows below are examples to adapt.Pre-launch checklist
Doing this in Fini
In Fini (usefini.com), the test plan above maps to these features:- Turn real conversations into test cases in Test Suite. Each case replays the customer’s messages, generates new replies, and evaluates them against your criteria: an AI judgement for outcomes like “grounded in the refund policy”, and exact checks for the Action, rule, article, reply, or handoff you expect. Mark the criteria that must hold Required to pass.
- Share one standard across many cases with Criteria groups (for example, one group for escalation cases), and group a launch scope into a collection you run with Run collection.
- Test drafts before publishing: in the run dialog, choose the Behavior version and select rule and article drafts. Nothing is published by a run.
- Test Suite uses recorded or mock Action responses, never live calls. Override recorded Action responses in Test setup to cover both success and failure outcomes, and do live end-to-end Action checks in a sandbox environment or a controlled pilot.
- Explore single edge cases, read the AI Steps trace and replay fixes in Inbox, then lock each failure in with the flask icon, which creates a test case from that conversation.
- Diagnose a wrong reply and draft the fix with Refine with AI.
- Add reply checks such as Banned terms, Confidential attributes and URL allowlist in Guardrails, and test them before enabling them on production channels.
- Follow the Fini-specific release gate in Best practices for testing your agent.
Related
Best practices for testing your agent
The Fini release gate and how to organize test cases.
Test Suite
Test cases, criteria groups, collections and runs.
How to run an AI support pilot
Scope, baseline, success criteria and the go/no-go readout.
Guardrails
Reply checks that run before delivery.

