Skip to main content
Test an AI support agent before launch with a fixed group of test cases built from your own real tickets, a written expected outcome for each case, and grading criteria that check whether every policy answer traces back to an approved source. Then rerun those same cases after every change to prompts, knowledge, workflows or models, so regressions show up in testing instead of in front of customers. This playbook covers how to build the cases, what to put in them, how to grade it, and how to decide when the agent is ready. It applies to any AI support agent, whatever platform you use.

Start from real tickets, not invented questions

Questions your team writes from memory are cleaner than what customers actually send. Pull a sample of recent closed tickets from your helpdesk and build the test cases from them, keeping the customer’s original wording (with personal data removed). Choose cases along three axes so they reflect your real workload and your real exposure: A low-volume, high-risk intent (for example, a complaint that may need formal handling) deserves cases even if it appears rarely. Volume tells you where most conversations land; risk tells you where a single failure hurts.

Write golden answers as outcomes, not scripts

A golden answer is the reference for what a correct response looks like. Write it as the outcome you expect across the conversation, not the exact words. AI agents phrase things differently each run, so word-for-word expectations fail for reasons that don’t matter. A useful golden answer has four parts:
  1. The facts that must appear. The refund window, the fee amount, the eligibility condition.
  2. The source those facts come from. The specific approved article, policy or system field.
  3. The action, if any. Which operation should run, with which inputs.
  4. The end state. Resolved, waiting for the customer, or handed to a human with context.
Then define grading criteria the whole team applies the same way. Pass/fail per criterion is easier to agree on than a single score. If you use an automated grader (an AI model that evaluates replies against written criteria), calibrate it against human grading on a sample first, and keep a human review of a share of results during the pilot period.

Cover the edge cases that break agents

Happy paths prove the agent works. Edge cases show you where it fails. Include each of these:

Catch wrong policy answers

The most damaging failure is a confident, fluent answer that states the wrong policy. It reads fine, so casual review misses it. Test for it on purpose.
  • Require traceability. For every policy case, the grader checks that each claim comes from an approved source. An answer that is correct by luck but not grounded is a fail, because it will be wrong next time.
  • Test conflicting-policy cases. If two documents disagree (an old help article and a newer internal policy, or two regions with different rules), write cases that hit the conflict. The expected outcome is that the agent uses the source you designated as authoritative, or escalates. Fix the conflict in the knowledge itself afterwards.
  • Test stale content. Include a case for any policy that changed recently. If the agent quotes the old version, find and retire the old source.
  • Test the gap. Ask something your knowledge doesn’t cover. The right answer is “I don’t have that information” plus a handoff, not a plausible guess.

Test actions in a sandbox

If the agent can take actions (issue refunds, change plans, update addresses), test them against a sandbox or staging environment, never production. Check that:
  • The right action runs only when its conditions are met, and doesn’t run when they aren’t.
  • Inputs are correct: the right account, amount and item.
  • Failures are handled. Make the sandbox return an error or a timeout and confirm the agent tells the customer accurately and hands off.
  • Confirmation is asked for where your policy requires it, before anything irreversible.

Test escalation both ways

Write cases that must escalate (explicit request for a human, legal threat, vulnerable customer, compliance trigger) and cases that must not (a customer venting about a delay the agent can resolve). For each handoff, check that the human receives enough context to continue without asking the customer to repeat themselves.

Run the loop on every change

Testing is not a one-time gate. Every change to a prompt, knowledge source, workflow, action or model can break something that used to work. Rerun the cases each time and compare against previous runs. Two habits keep the loop useful. First, change one thing at a time, so a drop in results points to one cause. Second, add every real failure you find, in testing or in production, as a new case. Your cases then grow with what you learn.

Set pass thresholds as a team decision

There is no universal score that means “ready”. Agree on thresholds before the first run, with support, compliance and product in the room, and write them down. A common pattern is to set them per category rather than overall:
  • Zero tolerance categories, such as wrong policy statements on regulated topics, unauthorized actions or data shown to the wrong person. Any failure blocks launch.
  • High bar categories, such as top-volume intents and escalation cases.
  • Improving categories, such as tone or rare intents, where you accept some misses and track the trend.
Deciding this in advance stops the threshold from drifting to whatever the latest run happened to score.

Sample test-case table

Use a simple table (a spreadsheet works) as the source of truth for your cases. The rows below are examples to adapt.

Pre-launch checklist

Doing this in Fini

In Fini (usefini.com), the test plan above maps to these features:
  • Turn real conversations into test cases in Test Suite. Each case replays the customer’s messages, generates new replies, and evaluates them against your criteria: an AI judgement for outcomes like “grounded in the refund policy”, and exact checks for the Action, rule, article, reply, or handoff you expect. Mark the criteria that must hold Required to pass.
  • Share one standard across many cases with Criteria groups (for example, one group for escalation cases), and group a launch scope into a collection you run with Run collection.
  • Test drafts before publishing: in the run dialog, choose the Behavior version and select rule and article drafts. Nothing is published by a run.
  • Test Suite uses recorded or mock Action responses, never live calls. Override recorded Action responses in Test setup to cover both success and failure outcomes, and do live end-to-end Action checks in a sandbox environment or a controlled pilot.
  • Explore single edge cases, read the AI Steps trace and replay fixes in Inbox, then lock each failure in with the flask icon, which creates a test case from that conversation.
  • Diagnose a wrong reply and draft the fix with Refine with AI.
  • Add reply checks such as Banned terms, Confidential attributes and URL allowlist in Guardrails, and test them before enabling them on production channels.
  • Follow the Fini-specific release gate in Best practices for testing your agent.

Best practices for testing your agent

The Fini release gate and how to organize test cases.

Test Suite

Test cases, criteria groups, collections and runs.

How to run an AI support pilot

Scope, baseline, success criteria and the go/no-go readout.

Guardrails

Reply checks that run before delivery.