> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usefini.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Best practices for testing your agent

> A practical testing playbook for Fini: which surface to use for what, what to cover, how to organize test cases, criteria groups and collections, and the regression check to run before every change ships.

export const ChecklistMeter = ({title = "Checklist", items = [], results = {}, disclaimer}) => {
  const FV = {
    lime: "#C3EE5E",
    ink: "#131415",
    line: "rgba(127,127,127,0.28)",
    soft: "rgba(127,127,127,0.07)",
    softer: "rgba(127,127,127,0.04)",
    muted: "rgba(127,127,127,0.95)",
    pass: "#C3EE5E",
    warn: "#FFB020",
    fail: "#FF4D4D",
    radius: 14
  };
  const fvCard = {
    border: `1px solid ${FV.line}`,
    borderRadius: FV.radius,
    padding: 18,
    margin: "20px 0",
    background: FV.softer
  };
  const fvChip = active => ({
    border: `1px solid ${active ? FV.lime : FV.line}`,
    background: active ? FV.lime : "transparent",
    color: active ? FV.ink : "inherit",
    borderRadius: 999,
    padding: "6px 12px",
    fontSize: 13,
    fontWeight: 600,
    cursor: "pointer",
    lineHeight: 1.2
  });
  const fvBtn = primary => ({
    border: `1px solid ${primary ? FV.lime : FV.line}`,
    background: primary ? FV.lime : "transparent",
    color: primary ? FV.ink : "inherit",
    borderRadius: 10,
    padding: "7px 14px",
    fontSize: 13,
    fontWeight: 600,
    cursor: "pointer"
  });
  const fvLabel = {
    fontSize: 11,
    fontWeight: 700,
    letterSpacing: "0.08em",
    textTransform: "uppercase",
    opacity: 0.6,
    marginBottom: 8
  };
  const [on, setOn] = useState(() => items.map(() => false));
  const n = on.filter(Boolean).length;
  const reqMissing = items.some((it, i) => it.required && !on[i]);
  const pct = items.length ? Math.round(n / items.length * 100) : 0;
  const msg = n === items.length ? results.complete : reqMissing && results.missingRequired ? results.missingRequired : results.partial;
  return <div style={fvCard}>
      <div style={{
    display: "flex",
    justifyContent: "space-between",
    alignItems: "baseline"
  }}>
        <div style={fvLabel}>{title}</div>
        <div style={{
    fontSize: 13,
    fontWeight: 700
  }}>{n} / {items.length}</div>
      </div>
      <div style={{
    height: 8,
    borderRadius: 999,
    background: FV.soft,
    overflow: "hidden",
    marginBottom: 12
  }}>
        <div style={{
    width: `${pct}%`,
    height: "100%",
    background: FV.lime,
    transition: "width .3s"
  }} />
      </div>
      {items.map((it, i) => <label key={i} style={{
    display: "flex",
    gap: 10,
    alignItems: "flex-start",
    padding: "8px 4px",
    borderTop: i ? `1px solid ${FV.line}` : "none",
    cursor: "pointer"
  }}>
          <input type="checkbox" checked={on[i]} onChange={() => setOn(o => o.map((v, j) => j === i ? !v : v))} style={{
    marginTop: 3,
    accentColor: FV.lime
  }} />
          <span style={{
    fontSize: 14,
    lineHeight: 1.5
  }}>
            {it.label}{it.required && <span style={{
    fontSize: 11,
    fontWeight: 700,
    marginLeft: 6,
    opacity: 0.6
  }}>REQUIRED</span>}
            {it.detail && <span style={{
    display: "block",
    fontSize: 12.5,
    opacity: 0.65
  }}>{it.detail}</span>}
          </span>
        </label>)}
      {msg && <div style={{
    marginTop: 12,
    fontSize: 13.5,
    padding: "10px 12px",
    borderRadius: 10,
    background: n === items.length ? "rgba(195,238,94,0.16)" : FV.soft
  }}>{msg}</div>}
      {disclaimer && <div style={{
    fontSize: 12,
    opacity: 0.6,
    marginTop: 8
  }}>{disclaimer}</div>}
    </div>;
};

Test your Fini (usefini.com) agent in layers: **Suggested scenarios** and **▶ Test Run** for each Rulebook rule, **Inbox** conversations and replays to verify individual fixes, and **Test Suite** test cases as the regression check you run before every change ships. Each layer catches a different kind of failure, and together they let you change prompts, knowledge, and rules without customers finding the regressions first.

This page is the playbook: which tool to use when, what to test, how to organize test cases, criteria groups and collections, and how to run them safely.

## Pick the right tool for the job

| Tool | Best for | Scope | Where |
| - | - | - | - |
| **Suggested scenarios** | Branch coverage for one intent rule, before it's published | One rule, every likely path | [Rulebook → Testing a rule](/en/automations/rulebook#suggested-scenarios) |
| **▶ Test Run** | Running one rule against a real conversation link or custom JSON | One rule, one input | [Rulebook → Manual test run](/en/automations/rulebook#manual-test-run) |
| **New conversation in Inbox** | Exploring an edge case by hand, with custom metadata | One conversation, the whole agent | [Inbox → Testing new conversations](/en/testing/inbox#testing-new-conversations) |
| **Replay** | Confirming a fix on a known-bad conversation | One conversation, re-run against a selected agent version | [Inbox → Replays](/en/testing/inbox#replays-and-regeneration), [Replays API](/en/api-reference/replays) |
| **Refine with AI** | Diagnosing a wrong reply and drafting the fix | One reply | [Refine with AI](/en/testing/fix-with-ai) |
| **Test Suite** | Regression across many conversation-based test cases, each evaluated against its criteria | A case, a collection, or **All tests** | [Test Suite](/en/testing/test-suite) |

The pattern is the one [Inbox](/en/testing/inbox#how-inbox-fits) is built around: find a failure, trace it in **AI Steps**, fix it at the source, replay to confirm, then lock it in as a Test Suite test case.

```mermaid theme={null}
---
title: The test loop
---
flowchart LR
    FAIL(("Failure found<br/>Inbox, feedback, Analytics"))
    TRACE["Trace it<br/>AI Steps"]
    FIX["Fix at the source<br/>Refine with AI"]
    REPLAY["Replay<br/>confirm the fix"]
    LOCK["Lock it in<br/>Test Suite test case"]
    GATE["Regression check<br/>run the collection"]
    SHIP["Publish"]

    FAIL --> TRACE --> FIX --> REPLAY --> LOCK --> GATE --> SHIP
    GATE -.->|"required criterion fails"| FIX
    SHIP -.->|"next failure"| FAIL

    classDef source fill:#F7F7F7,color:#131415,stroke:#E8E8E8
    classDef agent fill:#131415,color:#FFFFFF,stroke:#131415,stroke-width:3px
    classDef surface fill:#FFFFFF,color:#131415,stroke:#131415
    classDef human fill:#C3EE5E,color:#131415,stroke:#131415,stroke-width:2px

    class FAIL agent
    class TRACE,FIX,REPLAY surface
    class LOCK,GATE human
    class SHIP source
```

## How Test Suite works

Every test case in [Test Suite](/en/testing/test-suite) starts from a real conversation. A run replays the customer's messages from that conversation, generates new agent replies, and evaluates them against the case's criteria. Test Suite does not invent new customer scenarios: to test another customer path, create or select another conversation.

| Building block | What it is | Use it for |
| - | - | - |
| **Test case** | One conversation plus its **Test setup** (recorded Action responses and conversation attributes) and its **Criteria**. | One behavior you want to keep working. |
| **Criteria** | Checks under **What do you want to check?**: an AI judgement of a written expectation, an exact check for an Action, rule, article, reply, or handoff, a reusable criterion, or an advanced deterministic condition. | Defining what a passing result means. Mark the ones that decide the verdict **Required to pass**. |
| **Criteria groups** | Shared criteria linked to many cases. | One standard (for example "hands off on an explicit request for a human") applied across many conversations. |
| **Collections** | A named set of cases and criteria groups. | Scoping a run to one feature or release. |
| **Runs** | Saved results for every case run. | Comparing verdicts over time. |

<Info>
  Test Suite uses recorded or mock responses for external Actions. There is no live Action execution in Test Suite, and running with draft Behavior, rules, or articles does not publish those drafts. For live end-to-end checks of an Action, use a sandbox environment or a controlled [pilot](/en/playbooks/running-a-pilot).
</Info>

## What to test

Cover each of these categories for every agent. Most teams start with the first three and add the rest within the first weeks.

| Category | What you're checking | Suggested check or tool |
| - | - | - |
| **Happy paths** | The agent answers your highest-volume questions correctly and completes your main workflows. | AI judgement of a written expectation, plus an exact article check where one article should always be used |
| **Paraphrases** | The agent matches intent, not keywords: "I forgot my login" finds the password reset article. | Exact article check, linked across several conversations through one criteria group |
| **Edge cases** | Missing data, ineligible customers, two orders instead of one, a null User Attribute. | Rulebook **Suggested scenarios**, **▶ Test Run** with **Custom JSON**, and Test Suite cases with overridden attributes in **Test setup** |
| **Escalation** | The agent hands off when it should (explicit request, compliance concern) and doesn't when it shouldn't (a customer venting). | Exact handoff check, **Required to pass** |
| **Actions** | The right [Action](/en/api-reference/actions) runs at the right time, and failures are handled. | Exact Action check, with **Success** and **Failure** mock responses in **Test setup** |
| **Guardrails and scope** | The agent declines out-of-scope requests, doesn't leak confidential attributes, and respects your policies. | AI judgement written as the boundary you expect, plus your [Guardrails](/en/configuration/guardrails) |
| **Tone** | Replies sound like your brand and match the customer's emotional register. | AI judgement, usually not **Required to pass** |
| **Multilingual** | The agent answers correctly in the languages your customers use, from the right language folders. | Cases built from real conversations in each language, with an AI judgement and an exact article check |

<Tip>
  Guardrail channel scope does not exclude dashboard tests, Test Suite, or replay evaluation. That means you can test a new guardrail policy against your test cases before you enable it on a production channel.
</Tip>

Use an AI judgement when the outcome needs interpretation, such as whether a reply explains a refund policy without inventing an exception. Use an exact check when the outcome is an observable event, such as whether the agent called a particular Action or handed off to a human. Exact checks don't vary between runs, so lean on them for anything that must never break.

## Organize cases, groups, and collections

Keep each criteria group focused on one behavior. A failing run then tells you immediately what regressed. A small, well-chosen set of cases beats a large random one.

| Criteria group | Shared criteria | Conversations to link | Size to start | Run when |
| - | - | - | - | - |
| `Golden conversations` | AI judgement: "Resolves the request the way the original reply did, without inventing policy" | Hand-picked successes from Inbox | 10 to 40 | Every change |
| `Top intents: knowledge` | Exact article check plus an AI judgement on accuracy | Real questions for each top intent | 15 to 30 | Knowledge or prompt changes |
| `Escalation behavior` | Exact handoff check, **Required to pass** | Conversations filtered to escalation tags in Inbox. Keep cases that should not escalate in a separate group with the opposite check | 10 to 20 | Rulebook, prompt, or Reply Rules changes |
| `Refunds workflow` (one per critical workflow) | Exact Action and rule checks | Conversations that ran the workflow, with mock **Success** and **Failure** responses | 8 to 15 | Changes to that rule or its Actions |
| `Safety and scope` | AI judgement written as the boundary | Conversations you create in Inbox with out-of-scope and adversarial requests | 10 to 20 | Prompt, guardrail, or model changes |
| `Spanish` (one per major language) | AI judgement plus exact article check | Real conversations in that language | 10 to 20 | Knowledge, translation, or prompt changes |
| `Fixed bugs` | Case-specific criteria written for each fix | Each conversation you fix, added with the flask icon in Inbox | Grows over time | Every change |

A criteria group's conversation picker can add up to 50 conversations at a time from Inbox. For a single conversation, use the flask icon in Inbox, select it, and click **Select** to open the case editor.

Then group by release scope: put the cases and criteria groups for one feature or launch into a **collection**, and use **Run collection** before that feature changes. Use **All tests** for the full regression pass.

Write each AI judgement as the outcome you want across the whole conversation, not exact wording. For example: *"Agent verifies the account, explains the refund window, and escalates to billing only if the order is outside it."*

<Note>
  Editing a criteria group's criteria changes what its linked cases use on their next run. Saved run results do not change. Review the affected cases before confirming a group update.
</Note>

## Run the regression check before every change

Make this the release gate for any change to a prompt, knowledge source, Rulebook rule, Action, or model.

<Steps>
  <Step title="Test the change in isolation">
    For a rule change, run **Suggested scenarios** on the draft. For a knowledge or prompt change, open the conversations it should affect in Inbox and replay them against the version you're testing.
  </Step>

  <Step title="Run the affected cases against your drafts">
    You don't need to publish first. In the run dialog, under **Behaviour**, choose **Latest published behaviour** or a saved version, and select the rule and article drafts you want to test. Selected drafts replace their published versions for that run only. Nothing is published. Use **Run collection** for the affected scope or **Run all** for everything.
  </Step>

  <Step title="Read the verdicts, not just the count">
    Open **Runs**. Separate **Failed** (a required criterion did not pass) from **Execution error** (the run could not complete normally). Expand each failed result and read the criterion-level reasoning, the generated conversation, and the Action responses.
  </Step>

  <Step title="Fix, then re-run">
    Fix at the source, then use **Rerun test case** or run the collection again. If an execution error reports a missing Action response, click **Add mock response**, save it, and rerun.
  </Step>

  <Step title="Publish the change">
    Once the required criteria pass, publish the prompt, rule, or article drafts you tested.
  </Step>

  <Step title="Watch production after release">
    Check [Analytics](/en/analytics) the following days for changes in AI resolution rate, escalation reasons, and CSAT. See [Measure your resolution rate](/en/how-to/measure-resolution-rate#compare-before-and-after-a-change).
  </Step>
</Steps>

<Warning>
  Test Suite never calls your live Action endpoints: a missing response produces an execution error, not a live call. Two other surfaces can reach your real systems: creating a new conversation in Inbox may call live integrations, and an Inbox replay with **Run live actions** runs external actions again. If a conversation cancels a subscription or updates an address, it actually happens. Point Actions at sandbox URLs for any agent you use for live end-to-end testing.
</Warning>

<ChecklistMeter
  title="Before this change ships"
  items={[
{ label: "Tested the change in isolation", detail: "Suggested scenarios for a rule change; a replay of the affected conversations for a knowledge or prompt change.", required: true },
{ label: "Ran the affected collection with the drafts selected", detail: "Under Behaviour, pick the version to test and select the rule and article drafts. Nothing is published.", required: true },
{ label: "Every required criterion passes", detail: "Compare with earlier runs. A new failure without an intentional criteria change is a regression.", required: true },
{ label: "Read the reasoning on every Failed result, fixed, and re-ran", required: true },
{ label: "Execution errors resolved", detail: "Missing Action responses added with Add mock response, so every case actually ran.", required: true },
{ label: "Live Action checks done only in a sandbox or a controlled pilot", detail: "Test Suite uses recorded or mock Action responses, so it cannot prove a live endpoint works.", required: true },
{ label: "Changed one variable at a time", detail: "So a changed verdict points to one cause." },
{ label: "Planned to watch Analytics after release", detail: "AI resolution rate, escalation reasons, and CSAT over the following days." }
]}
  results={{
complete: "Ready to publish. Keep an eye on Analytics over the next few days.",
partial: "The release gate is met. The remaining habits make the next comparison cleaner.",
missingRequired: "Not ready yet. Finish the required checks before publishing."
}}
/>

## Habits that keep tests useful

* **Add every fixed bug as a test case.** Before you close the loop on a wrong reply, use the flask icon in Inbox to create a case from that conversation and write the criterion that would have caught it. Your suite then grows with every failure you learn about.
* **Change one variable at a time.** If you edit a prompt and a rule together, a changed verdict can't tell you which one caused it.
* **Check the run configuration before blaming the agent.** Compare the recorded criteria, test setup, and selected draft versions with the earlier run. A changed verdict can come from an edited criterion or setup rather than a different agent.
* **Write AI judgements as outcomes.** Expectations that describe the outcome you want are stable across runs; expectations that demand exact wording fail on harmless rephrasing.
* **Separate errors from failures.** An execution error means the case could not complete normally, often because a recorded Action response is missing. Fix the setup before you read the verdict.
* **Test both Action outcomes.** Use **Duplicate test case** to copy a case, then set one copy's mock response to **Success** and the other to **Failure**. The copy's setup and case-specific criteria are independent; linked group criteria stay shared.
* **Keep required checks few and sharp.** Mark only the criteria that define the verdict as **Required to pass**. Non-required checks still show in results without failing the case.

## Automate parts of it with the API

Test Suite runs are started from the dashboard. The public API for the rebuilt Test Suite is not documented yet, so build CI checks around the endpoints below.

| Task | Endpoint |
| - | - |
| Replay one conversation turn against a selected agent configuration | `POST /v2/replays/public`, with `ruleExecutionMode` defaulting to `simulate` |
| Generate and run deterministic paths for a rule version | `GET /v2/hc-rules/{id}/versions/{versionId}/test-paths/public` and `POST /v2/hc-rules/{id}/versions/{versionId}/test-paths/run/public` |
| Draft a fix for a wrong reply | [Create fix-review iteration](/en/api-reference/create-fix-review-iteration) |

Replays create separate replay conversations. They don't overwrite the original conversation or apply any changes by themselves. See [Replays](/en/api-reference/replays) for modes and model overrides, and [API Keys](/en/deploy/api-keys) for scopes.

## Related

<CardGroup cols={2}>
  <Card title="Test Suite" icon="vial" href="/en/testing/test-suite">
    Test cases, criteria groups, collections, and runs.
  </Card>

  <Card title="Inbox" icon="inbox" href="/en/testing/inbox">
    AI Steps, replays, feedback, and creating test cases from conversations.
  </Card>

  <Card title="Refine with AI" icon="wand-magic-sparkles" href="/en/testing/fix-with-ai">
    Turn a wrong reply into a reviewed fix.
  </Card>

  <Card title="Intent Rules" icon="diagram-project" href="/en/automations/rulebook">
    Suggested scenarios and manual test runs for each rule.
  </Card>
</CardGroup>


## Related topics

- [Testing an AI support agent before launch](/en/playbooks/testing-before-launch.md)
- [Questions to ask an AI support vendor, with Fini's answers](/en/evaluate/questions-to-ask.md)
- [Changelog](/en/changelog.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.