Choose the scope
A pilot answers one question: does the agent resolve this kind of work well enough to expand? A broad scope makes that question impossible to answer. Pick a scope you can describe in one sentence and write it down.
Also name the owners: who reviews conversations, who fixes knowledge, who can change the agent’s configuration, and who signs off at the end.
Capture a baseline before you start
Without a baseline, every result is unfalsifiable. Before the agent goes live, pull the same metrics you’ll use to judge the pilot, for the same scope, over a recent comparable period:- Volume per intent and channel.
- Time to first response and time to resolution.
- CSAT for the in-scope intents.
- Reopen or repeat-contact rate, if your helpdesk tracks it.
- Current escalation or transfer patterns.
Define success criteria
Agree on the criteria and thresholds with everyone who will sign off, before day one. These four cover most pilots:
Add a quality check on top: a reviewer grades a sample of resolved conversations each week for correctness and grounding (each answer traceable to an approved source). A conversation the system marks resolved but that gave the wrong policy is a failure, and the metrics alone won’t show it.
Shadow mode, then live
Shadow mode means the agent drafts a reply for every in-scope conversation but only your team sees it, as an internal note. Customers are still answered by humans. It lets you measure answer quality on real traffic with no customer exposure. Live means the agent replies to customers directly.
Most pilots run shadow mode first, then go live intent by intent, keeping high-risk intents in shadow mode longer.
Get enough conversations to judge
You don’t need a precise statistical design, but you do need enough conversations that the result isn’t noise. A few principles hold regardless of your volume:- Judge per intent, not just overall. An overall rate can hide one intent that fails every time. Each in-scope intent needs enough conversations for reviewers to see its failure patterns.
- Fix the window in advance. Decide the pilot length and review dates before you start. Ending early because the numbers look good (or extending because they don’t) biases the result.
- Account for cycles. Weekly patterns, month-end billing and seasonal peaks change what customers ask. A window that covers at least one full cycle of your business is easier to trust.
- Don’t wait on rare, high-risk intents. If an intent is rare, you may never get enough live examples. Cover it with targeted test cases before launch instead.
- Keep changes visible. You will fix things during the pilot. Log every change with its date, so a jump in a metric can be tied to a cause.
Review weekly
Hold a short weekly review with the owners. Use the same agenda each time:- Metrics against baseline and against last week, per intent.
- A sample of resolved conversations, graded for correctness and grounding.
- A sample of handoffs, graded for timing and context.
- Failures found, their root cause (missing knowledge, conflicting knowledge, wrong workflow, wrong trigger) and the fix.
- Changes made since the last review, and changes planned.
- Scope decisions: move an intent from shadow to live, hold it, or pull it back.
The go/no-go readout
At the end of the window, write a one-page readout. Use the same structure every time, so decisions are comparable across pilots. The figures below are placeholders to replace with your own.Common failure modes
Doing this in Fini
In Fini (usefini.com), each pilot phase maps to these features:- Plan the phases with the Rollout timeline, which maps each milestone to the product steps and exit checks.
- Run shadow mode with an Internal Comment rule in Reply Rules, keep sensitive intents on Internal Comment or No Reply, and switch passed traffic to Direct Reply.
- Read your AI resolution rate in Analytics, slice it by channel, intent rule and escalation reason, and compare before and after each change with Measure your resolution rate.
- Build the pre-launch test cases in Test Suite from real conversations, group them into a collection for the pilot scope, and run it after every change, following Best practices for testing your agent.
- Test Suite uses recorded or mock Action responses, so it does not prove a live endpoint works. Verify Actions end to end in a sandbox environment first, then watch them closely on real pilot traffic. During shadow mode, point write Actions at a sandbox or keep the rules that call them unpublished until those intents go live.
- If you are a new enterprise customer, read 90-Day Money-Back Guarantee explained for the terms (eligibility applies).
Related
Rollout timeline
Day 1, 14 and 30 milestones with exit checks.
Measure your resolution rate
Read, slice and compare AI resolution rate.
Testing an AI support agent before launch
Build the test cases your pilot starts from.
Reply Rules
No Reply, Internal Comment and Direct Reply.

