> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usefini.com/llms.txt
> Use this file to discover all available pages before exploring further.

# How to run an AI support pilot

> How to scope an AI support pilot, capture a baseline, define success by resolution rate rather than deflection, move from shadow mode to live, and close with a go/no-go readout.

export const StepExplorer = ({title, steps = [], hint = "Click a step, or use Next."}) => {
  const FV = {
    lime: "#C3EE5E",
    ink: "#131415",
    line: "rgba(127,127,127,0.28)",
    soft: "rgba(127,127,127,0.07)",
    softer: "rgba(127,127,127,0.04)",
    muted: "rgba(127,127,127,0.95)",
    pass: "#C3EE5E",
    warn: "#FFB020",
    fail: "#FF4D4D",
    radius: 14
  };
  const fvCard = {
    border: `1px solid ${FV.line}`,
    borderRadius: FV.radius,
    padding: 18,
    margin: "20px 0",
    background: FV.softer
  };
  const fvChip = active => ({
    border: `1px solid ${active ? FV.lime : FV.line}`,
    background: active ? FV.lime : "transparent",
    color: active ? FV.ink : "inherit",
    borderRadius: 999,
    padding: "6px 12px",
    fontSize: 13,
    fontWeight: 600,
    cursor: "pointer",
    lineHeight: 1.2
  });
  const fvBtn = primary => ({
    border: `1px solid ${primary ? FV.lime : FV.line}`,
    background: primary ? FV.lime : "transparent",
    color: primary ? FV.ink : "inherit",
    borderRadius: 10,
    padding: "7px 14px",
    fontSize: 13,
    fontWeight: 600,
    cursor: "pointer"
  });
  const fvLabel = {
    fontSize: 11,
    fontWeight: 700,
    letterSpacing: "0.08em",
    textTransform: "uppercase",
    opacity: 0.6,
    marginBottom: 8
  };
  const [i, setI] = useState(0);
  const s = steps[i] || ({});
  return <div style={fvCard}>
      {title && <div style={fvLabel}>{title}</div>}
      <div style={{
    display: "flex",
    alignItems: "center",
    gap: 0,
    overflowX: "auto",
    paddingBottom: 6
  }}>
        {steps.map((st, k) => <div key={k} style={{
    display: "flex",
    alignItems: "center",
    flex: k === steps.length - 1 ? "0 0 auto" : "1 0 auto"
  }}>
            <button onClick={() => setI(k)} title={st.title} style={{
    width: 34,
    height: 34,
    borderRadius: 999,
    flex: "0 0 auto",
    cursor: "pointer",
    fontWeight: 700,
    fontSize: 13,
    border: `2px solid ${k <= i ? FV.lime : FV.line}`,
    background: k === i ? FV.lime : k < i ? "rgba(195,238,94,0.25)" : "transparent",
    color: k === i ? FV.ink : "inherit"
  }}>{k + 1}</button>
            {k < steps.length - 1 && <div style={{
    height: 2,
    minWidth: 18,
    flex: 1,
    background: k < i ? FV.lime : FV.line,
    margin: "0 4px"
  }} />}
          </div>)}
      </div>
      <div style={{
    display: "flex",
    gap: 6,
    flexWrap: "wrap",
    margin: "6px 0 14px"
  }}>
        {steps.map((st, k) => <span key={k} onClick={() => setI(k)} style={{
    fontSize: 12,
    cursor: "pointer",
    opacity: k === i ? 1 : 0.55,
    fontWeight: k === i ? 700 : 500,
    marginRight: 8
  }}>{st.label}</span>)}
      </div>
      <div style={{
    border: `1px solid ${FV.line}`,
    borderRadius: 12,
    padding: 16,
    background: FV.soft,
    minHeight: 120
  }}>
        {s.meta && <div style={{
    ...fvLabel,
    opacity: 0.75
  }}>{s.meta}</div>}
        <div style={{
    fontSize: 17,
    fontWeight: 700,
    marginBottom: 6
  }}>{s.title}</div>
        {s.body && <div style={{
    fontSize: 14.5,
    lineHeight: 1.6,
    opacity: 0.9
  }}>{s.body}</div>}
        {s.points && <ul style={{
    margin: "10px 0 0",
    paddingLeft: 18,
    fontSize: 14,
    lineHeight: 1.6
  }}>
            {s.points.map((p, k) => <li key={k}>{p}</li>)}
          </ul>}
      </div>
      <div style={{
    display: "flex",
    justifyContent: "space-between",
    alignItems: "center",
    marginTop: 12,
    gap: 8
  }}>
        <span style={{
    fontSize: 12,
    opacity: 0.6
  }}>{hint}</span>
        <div style={{
    display: "flex",
    gap: 8
  }}>
          <button style={fvBtn(false)} disabled={i === 0} onClick={() => setI(v => Math.max(0, v - 1))}>Back</button>
          <button style={fvBtn(true)} onClick={() => setI(v => (v + 1) % Math.max(1, steps.length))}>{i === steps.length - 1 ? "Start over" : "Next"}</button>
        </div>
      </div>
    </div>;
};

Run an AI support pilot on a narrow, written scope (a few channels, intents and customer segments), measure a baseline before the agent touches anything, and judge the result on resolution rate, CSAT, escalation quality and response time rather than deflection. Start in shadow mode, go live in stages, review weekly, and end with a go/no-go readout against criteria you agreed before day one.

This playbook walks through each of those decisions. It applies to any AI support agent and is written for the support, CX and engineering leads who own the pilot.

## Choose the scope

A pilot answers one question: does the agent resolve this kind of work well enough to expand? A broad scope makes that question impossible to answer. Pick a scope you can describe in one sentence and write it down.

| Dimension | Good pilot choice | Why |
| - | - | - |
| **Channels** | One or two channels, usually chat or email first | Fewer moving parts, and written channels are easier to review than voice. |
| **Intents** | A handful of high-volume, well-documented intents, plus one or two that need an action | Volume gives you enough conversations to judge; one action-based intent tests more than answering. |
| **Segments** | A defined slice of customers: a region, a plan, a language, or a share of traffic | Lets you compare against customers outside the pilot and limits exposure. |
| **Exclusions** | Intents and segments the agent must not handle yet | High-risk or regulated topics stay with humans until the agent has earned them. |

Also name the owners: who reviews conversations, who fixes knowledge, who can change the agent's configuration, and who signs off at the end.

## Capture a baseline before you start

Without a baseline, every result is unfalsifiable. Before the agent goes live, pull the same metrics you'll use to judge the pilot, for the same scope, over a recent comparable period:

* Volume per intent and channel.
* Time to first response and time to resolution.
* CSAT for the in-scope intents.
* Reopen or repeat-contact rate, if your helpdesk tracks it.
* Current escalation or transfer patterns.

Use the same definitions before and after. If your helpdesk measures first response in business hours today, measure it the same way during the pilot.

## Define success criteria

Agree on the criteria and thresholds with everyone who will sign off, before day one. These four cover most pilots:

| Criterion | What to measure | Watch out for |
| - | - | - |
| **Resolution rate** | Conversations the agent resolved end to end with no human handover, as a share of in-scope conversations | Don't substitute **deflection**. Deflection usually counts any conversation that didn't reach a human, including customers who gave up or never replied. It can look healthy while customers go unhelped. |
| **CSAT** | Customer ratings on AI-handled conversations, compared with your human baseline for the same intents | Response rates on surveys can differ between AI and human conversations, so compare like with like. |
| **Escalation quality** | Did the agent hand off when it should, and did the human get enough context to continue? | Sample handoffs and grade them. A low escalation rate is not good if the wrong conversations stayed with the agent. |
| **Time to first response** | Time from the customer's first message to the first useful reply | Fast and wrong is worse than slow and right. Read this alongside resolution and accuracy. |

Add a quality check on top: a reviewer grades a sample of resolved conversations each week for correctness and grounding (each answer traceable to an approved source). A conversation the system marks resolved but that gave the wrong policy is a failure, and the metrics alone won't show it.

<Warning>
  Write down how you'll handle a serious error (a wrong policy statement on a regulated topic, an action on the wrong account) before the pilot starts: who is told, how fast, and whether the agent is paused. Check your regulator's rules for anything that applies to your industry.
</Warning>

## Shadow mode, then live

**Shadow mode** means the agent drafts a reply for every in-scope conversation but only your team sees it, as an internal note. Customers are still answered by humans. It lets you measure answer quality on real traffic with no customer exposure.

**Live** means the agent replies to customers directly.

| | Shadow mode | Live |
| - | - | - |
| **Customer exposure** | None | Full, within scope |
| **What you learn** | Answer correctness, grounding, coverage gaps, whether escalation triggers fire | Real resolution rate, CSAT, response time, customer behavior |
| **What you can't learn** | Resolution rate and CSAT, because customers never see the reply | Nothing hidden, but mistakes reach customers |
| **Good for** | The first phase, and any high-risk intent | Intents that passed shadow review |

Most pilots run shadow mode first, then go live intent by intent, keeping high-risk intents in shadow mode longer.

## Get enough conversations to judge

You don't need a precise statistical design, but you do need enough conversations that the result isn't noise. A few principles hold regardless of your volume:

* **Judge per intent, not just overall.** An overall rate can hide one intent that fails every time. Each in-scope intent needs enough conversations for reviewers to see its failure patterns.
* **Fix the window in advance.** Decide the pilot length and review dates before you start. Ending early because the numbers look good (or extending because they don't) biases the result.
* **Account for cycles.** Weekly patterns, month-end billing and seasonal peaks change what customers ask. A window that covers at least one full cycle of your business is easier to trust.
* **Don't wait on rare, high-risk intents.** If an intent is rare, you may never get enough live examples. Cover it with targeted test cases before launch instead.
* **Keep changes visible.** You will fix things during the pilot. Log every change with its date, so a jump in a metric can be tied to a cause.

## Review weekly

Hold a short weekly review with the owners. Use the same agenda each time:

1. Metrics against baseline and against last week, per intent.
2. A sample of resolved conversations, graded for correctness and grounding.
3. A sample of handoffs, graded for timing and context.
4. Failures found, their root cause (missing knowledge, conflicting knowledge, wrong workflow, wrong trigger) and the fix.
5. Changes made since the last review, and changes planned.
6. Scope decisions: move an intent from shadow to live, hold it, or pull it back.

<StepExplorer
  title="Pilot phases"
  steps={[
{ label: "Phase 1", meta: "Plan", title: "Scope, baseline and criteria", body: "Write down the scope, owners and success criteria, and pull the baseline for the same scope before the agent goes live.", points: ["Channels, intents, segments and exclusions", "Baseline: volume, response time, CSAT, repeat contacts", "Thresholds agreed by everyone who signs off", "Serious-error plan written"] },
{ label: "Phase 2", meta: "Prepare", title: "Test before any traffic", body: "Build test cases from real tickets for the in-scope intents and run them until they meet your thresholds.", points: ["Golden answers with named sources", "Edge cases and escalation cases", "Actions tested in a sandbox"] },
{ label: "Phase 3", meta: "Shadow", title: "Run on real traffic, internally", body: "The agent drafts replies as internal notes. Reviewers grade them for correctness, grounding and escalation.", points: ["Grade a sample every week", "Fix knowledge gaps and conflicts at the source", "Promote intents that pass review"] },
{ label: "Phase 4", meta: "Live", title: "Go live intent by intent", body: "Passed intents reply to customers directly. High-risk intents stay in shadow mode longer.", points: ["Track resolution rate, CSAT, escalation quality, response time", "Log every change with its date", "Pull an intent back if quality drops"] },
{ label: "Phase 5", meta: "Decide", title: "Go/no-go readout", body: "Compare results with the baseline and the agreed criteria, then decide to expand, extend or stop.", points: ["Use the readout template", "Name what changes if you expand", "Record open risks and owners"] }
]}
/>

## The go/no-go readout

At the end of the window, write a one-page readout. Use the same structure every time, so decisions are comparable across pilots. The figures below are placeholders to replace with your own.

| Section | What to include | Example entry |
| - | - | - |
| **Scope** | Channels, intents, segments, dates, exclusions | Chat and email, 6 intents, EU customers, 6 weeks |
| **Baseline vs pilot** | Each success metric, before and during, per intent | Table of resolution rate, CSAT, first response, by intent |
| **Quality review** | Share of sampled conversations graded correct and grounded; serious errors and how they were handled | Sample size, pass count, list of serious errors |
| **Escalation review** | Handoff timing and context quality from the sample | Missed escalations, unnecessary escalations |
| **What changed** | Every fix made during the pilot, with dates | Knowledge conflicts resolved, workflow updated |
| **Decision** | Go, extend or stop, per intent | Go on 4 intents, extend 1, stop 1 |
| **Conditions** | What must be true to expand | Thresholds for the next scope, owners, review cadence |
| **Open risks** | Known gaps and who owns them | Rare intent untested live, owner named |

## Common failure modes

| Failure mode | What it looks like | How to avoid it |
| - | - | - |
| **No baseline** | "Response time is great" with nothing to compare against | Pull the baseline before the agent goes live. |
| **Measuring deflection** | A high number while CSAT drops and repeat contacts rise | Use resolution rate as the headline metric. |
| **Scope creep** | New intents added mid-pilot, so nothing is comparable | Hold new intents for the next phase. |
| **Messy knowledge** | The agent quotes outdated or conflicting articles | Audit and fix knowledge before and during shadow mode. |
| **No owner for fixes** | Failures are found every week and never fixed | Name who fixes knowledge, workflows and configuration. |
| **Moving thresholds** | Criteria change to fit the results | Agree and write them down before day one. |
| **Too short or too narrow** | Too few conversations per intent to see failures | Fix a window that covers a full business cycle. |
| **Skipping handoff review** | Low escalation looks good, but the wrong conversations stayed with the agent | Grade a sample of handoffs and non-handoffs weekly. |

## Doing this in Fini

In Fini (usefini.com), each pilot phase maps to these features:

* Plan the phases with the [Rollout timeline](/en/rollout-timeline), which maps each milestone to the product steps and exit checks.
* Run shadow mode with an **Internal Comment** rule in [Reply Rules](/en/automations/reply-behavior), keep sensitive intents on **Internal Comment** or **No Reply**, and switch passed traffic to **Direct Reply**.
* Read your **AI resolution rate** in Analytics, slice it by channel, intent rule and escalation reason, and compare before and after each change with [Measure your resolution rate](/en/how-to/measure-resolution-rate).
* Build the pre-launch test cases in [Test Suite](/en/testing/test-suite) from real conversations, group them into a collection for the pilot scope, and run it after every change, following [Best practices for testing your agent](/en/testing/best-practices).
* Test Suite uses recorded or mock Action responses, so it does not prove a live endpoint works. Verify Actions end to end in a sandbox environment first, then watch them closely on real pilot traffic. During shadow mode, point write Actions at a sandbox or keep the rules that call them unpublished until those intents go live.
* If you are a new enterprise customer, read [90-Day Money-Back Guarantee explained](/en/billing/money-back-guarantee) for the terms (eligibility applies).

## Related

<CardGroup cols={2}>
  <Card title="Rollout timeline" icon="calendar-days" href="/en/rollout-timeline">
    Day 1, 14 and 30 milestones with exit checks.
  </Card>

  <Card title="Measure your resolution rate" icon="chart-line" href="/en/how-to/measure-resolution-rate">
    Read, slice and compare AI resolution rate.
  </Card>

  <Card title="Testing an AI support agent before launch" icon="flask-vial" href="/en/playbooks/testing-before-launch">
    Build the test cases your pilot starts from.
  </Card>

  <Card title="Reply Rules" icon="turn-down-right" href="/en/automations/reply-behavior">
    No Reply, Internal Comment and Direct Reply.
  </Card>
</CardGroup>


## Related topics

- [Testing an AI support agent before launch](/en/playbooks/testing-before-launch.md)
- [Questions to ask an AI support vendor, with Fini's answers](/en/evaluate/questions-to-ask.md)
- [Running multilingual AI support well](/en/playbooks/multilingual-support.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.