> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usefini.com/llms.txt
> Use this file to discover all available pages before exploring further.

# How accuracy and resolution rate are measured

> What Fini's stated benchmark of a 90% resolution rate at 99% accuracy means, how the product computes AI resolution rate, and how you can assess accuracy on your own conversations.

Fini (usefini.com) states a benchmark of **a 90% resolution rate at 99% accuracy** across its deployments, with accuracy measured through human review of conversations. In the product, resolution is measured precisely: **AI resolution rate** is the share of conversations whose status is **Resolved by AI**, and you can check accuracy yourself on your own traffic with Inbox feedback, Test Suite run results, guardrail activity, and CSAT.

The two numbers answer different questions. Resolution rate tells you how much of your volume the agent closed without a human. Accuracy tells you whether what the agent said and did was correct. A high resolution rate with poor accuracy is a liability, especially in fintech, banking, insurance, and healthcare, so this page covers both and shows where each one lives in the product.

<Note>
  The benchmark is Fini's stated figure across deployments, not a guarantee for every account. Your own numbers depend on your knowledge coverage, the workflows you automate, the Actions you connect, and how much of your traffic needs a human by policy. The rest of this page shows you how to measure your own.
</Note>

## What the headline numbers mean

| Number | What it describes | Where you verify it for your account |
| - | - | - |
| **90% resolution rate** | The share of conversations the agent resolves end to end with no human handover. | **AI resolution rate** in [Analytics](/en/analytics), computed from the **Resolved by AI** status. |
| **99% accuracy** | Fini's stated benchmark for how often the agent's answers are correct, measured through human review of conversations. | Your own review signals: Inbox feedback, Test Suite run results, guardrail hits, and CSAT. See [Assessing accuracy yourself](#assessing-accuracy-yourself). |

## How resolution rate is computed in the product

Every conversation the agent handles ends in exactly one of three statuses. The status comes from the agent's **Output Tag Selection** on the conversation (the **Conversation Status** category), which you can inspect in the [AI Steps trace](/en/testing/inbox#ai-steps-what-the-agent-actually-did).

| Status | Meaning | Counts toward AI resolution rate? |
| - | - | - |
| **Resolved by AI** | The agent resolved the conversation without a human. | Yes |
| **Escalated to Human Team** | The agent handed the conversation to a teammate. | No |
| **Waiting for Customer** | The agent replied and is waiting on the customer. | No |

The three statuses sum to 100% of conversations in the selected window. From them, Analytics derives three rates:

| Metric | Formula | What it tells you |
| - | - | - |
| **AI resolution rate** | Resolved by AI conversations / total conversations | How much of your volume the agent actually closed. This is the metric Fini treats as the headline. |
| **Human escalation rate** | Escalated to Human Team conversations / total conversations | How much of your volume reached your team. |
| **Deflection rate** | 100% - Human escalation rate | How much of your volume did not reach a human. It counts **Waiting for Customer** conversations as deflected. |

Because deflection includes conversations that are still waiting on the customer, deflection rate is always greater than or equal to AI resolution rate. The gap between the two is your **Waiting for Customer** share. [Resolution vs deflection](/en/performance/resolution-vs-deflection) walks through a worked example.

Where you see each number:

* **KPI cards** at the top of [Analytics](/en/analytics#kpi-cards) show **Deflection rate** and **Human escalation rate** with period-over-period change pills.
* The **Resolution rate** trend view plots the daily resolution rate alongside the **Waiting for Customer** share.
* The **Conversation Status** doughnut shows the full three-way split.
* **Knowledge performance** and **Intent rule breakdown** show **AI Resolve Rate** and **Escalated Rate** per knowledge slice and per [Rulebook](/en/automations/rulebook) intent rule.
* The [Get agent analytics](/en/api-reference/get-agent-analytics) API returns `aiResolutionRate`, `humanEscalationRate`, and the raw counts (`resolvedConversations`, `escalatedConversations`, `waitingForCustomerConversations`, `totalConversations`).

For a step-by-step procedure, see [Measure your resolution rate](/en/how-to/measure-resolution-rate).

## Why a resolved conversation is not automatically an accurate one

**Resolved by AI** means the agent closed the conversation without a human. It does not, on its own, prove that every reply was correct. A customer can accept a wrong answer and leave. That is why Fini pairs resolution with accuracy signals you control, and why every number in Analytics traces back to individual conversations you can open in [Inbox](/en/testing/inbox) and audit reply by reply.

## Assessing accuracy yourself

You don't have to take any benchmark on trust. Fini gives you several independent ways to grade the agent's accuracy on your own conversations. Use more than one: each catches a different kind of error.

| Signal | Where it lives | What it measures | How to use it |
| - | - | - | - |
| **Teammate review in Inbox** | [Inbox](/en/testing/inbox) per-message actions | Whether a specific reply was right, as judged by your team. | Sample conversations each week, mark replies with thumbs up or thumbs down, and add a **Feedback note** explaining what was wrong. Filter by **Feedback** and **Feedback Notes** to track the share of reviewed replies marked wrong. |
| **Test Suite runs** | [Test Suite](/en/testing/test-suite) | Whether the agent still handles a fixed set of real conversations correctly. Each run replays the customer's messages, generates new replies, and evaluates them against your criteria. | Write criteria as AI judgements (for example, "states the refund window from the policy without inventing an exception") and exact checks for the expected Action, article, or handoff. Required criteria that keep passing across runs mean accuracy is holding as you change the agent. Read each **Failed** result's reasoning, and treat an **Execution error** as a setup problem, not a wrong answer. |
| **Refine with AI diagnoses** | [Refine with AI](/en/testing/fix-with-ai) | The root cause of a specific wrong reply (prompt, knowledge, or rule). | Use the **Summary** and **Why this change** output to categorize errors, then lock each fix in as a Test Suite test case. |
| **Guardrail hits** | [Guardrails](/en/configuration/guardrails) activity | How often a generated reply initially failed a policy check (banned terms, confidential attributes, URL allowlist, AI disclosure, custom rules). | Review **Firing** policies over **24h**, **7d**, or **30d**. Hits count initial failed checks, including replies that were successfully rewritten, so they are not a count of escalations. |
| **Escalation reasons** | [Analytics](/en/analytics#escalation-reasons) | Why the agent stopped and handed off: **Missing Knowledge**, **Conflicting Knowledge**, **API or System Failure**, and so on. | A rising **Conflicting Knowledge** or **Partially Available** share points at knowledge accuracy problems before they become wrong answers. |
| **CSAT** | Analytics **Average CSAT** card and **CSAT** filter | What customers explicitly said about the conversation. | Compare CSAT on AI-resolved conversations against escalated ones. |
| **AI CSAT and sentiment** | Analytics **AI CSAT** filter and **Sentiment** filter | Fini's model-inferred satisfaction and tone, available on conversations the customer never rated. | Read alongside explicit CSAT. A thumbs-up after a frustrating exchange (high CSAT, low sentiment) is worth reading. |

The signals work as one loop: production traffic produces evidence of errors, you diagnose and fix them, and Test Suite locks each fix in so it can't quietly come back.

```mermaid theme={null}
---
title: The accuracy measurement loop
---
flowchart LR
    CONV(("Live conversations"))
    subgraph SIGNALS["Signals you read"]
        direction TB
        INBOX["Inbox review<br/>thumbs and Feedback notes"]
        GUARD["Guardrail hits"]
        ESCR["Escalation reasons"]
        CSAT["CSAT, AI CSAT<br/>and sentiment"]
    end
    REFINE["Refine with AI<br/>diagnose and draft the fix"]
    FIX["Approve the fix<br/>prompt, knowledge or rule"]
    SUITE["Test Suite<br/>regression check"]
    AGENT["Updated agent"]

    CONV --> SIGNALS
    SIGNALS --> REFINE
    REFINE --> FIX
    FIX --> SUITE
    SUITE -->|"required criteria pass"| AGENT
    AGENT --> CONV

    classDef source fill:#F7F7F7,color:#131415,stroke:#E8E8E8
    classDef agent fill:#131415,color:#FFFFFF,stroke:#131415,stroke-width:3px
    classDef surface fill:#FFFFFF,color:#131415,stroke:#131415
    classDef human fill:#C3EE5E,color:#131415,stroke:#131415,stroke-width:2px

    class CONV,AGENT agent
    class INBOX,GUARD,ESCR,CSAT source
    class REFINE,FIX surface
    class SUITE human
```

### A simple accuracy audit you can repeat

<Steps>
  <Step title="Draw a sample">
    In [Inbox](/en/testing/inbox), select the agent and the last 7 days, set **Fini Touched** to *Yes*, and filter **Conversation status** to **Resolved by AI**. Pick a fixed number of conversations at random (for example, 50) rather than the ones you remember.
  </Step>

  <Step title="Grade every Fini reply">
    For each conversation, read the thread and open **AI Steps** on any reply you are unsure about. Mark each reply thumbs up or thumbs down. Add a **Feedback note** to every thumbs down that says what the correct answer was.
  </Step>

  <Step title="Compute your accuracy rate">
    Divide the replies you marked correct by the replies you graded. Record the result with the date range, the agent, and the sample size so the next audit is comparable.
  </Step>

  <Step title="Fix and lock in">
    Run [Refine with AI](/en/testing/fix-with-ai) on the wrong replies, approve the fixes you agree with, then use the flask icon in Inbox to create a [Test Suite](/en/testing/test-suite) test case from each conversation, so the error becomes a permanent regression check. Then mark the feedback as actioned, so the **Feedback Actioned** filter in Inbox shows which thumbs-downs are addressed.
  </Step>
</Steps>

<Tip>
  Grade escalated conversations too, with a different question: should the agent have escalated? Unnecessary escalations lower resolution rate; missed escalations are an accuracy and compliance problem. In Test Suite, link those conversations to a criteria group with an exact handoff check marked **Required to pass** to automate this check.
</Tip>

## What these numbers do not tell you

* **Analytics does not grade correctness.** It counts statuses, escalation reasons, and ratings. Correctness comes from your review, Test Suite, and guardrails.
* **Test Suite is a fixed sample.** Passing test cases mean the agent handles those conversations correctly. They do not cover traffic you haven't turned into test cases, which is why production review still matters. Test Suite also uses recorded or mock Action responses, so it does not show whether a live Action endpoint works.
* **Guardrails are not a fail-closed boundary.** A check can report **could not run** and the original reply can still be sent. See [Guardrails](/en/configuration/guardrails#reply-outcomes).

## Related

<CardGroup cols={2}>
  <Card title="Resolution vs deflection" icon="scale-balanced" href="/en/performance/resolution-vs-deflection">
    The exact definitions, a worked example, and why Fini leads with resolution.
  </Card>

  <Card title="Measure your resolution rate" icon="chart-line" href="/en/how-to/measure-resolution-rate">
    Step-by-step in Analytics and through the API, including before-and-after comparisons.
  </Card>

  <Card title="Analytics" icon="chart-bar" href="/en/analytics">
    Every KPI card, chart, and breakdown on the Analytics page.
  </Card>

  <Card title="Test Suite" icon="vial" href="/en/testing/test-suite">
    Conversation-based test cases evaluated against your criteria on every change.
  </Card>
</CardGroup>


## Related topics

- [RFP fact sheet](/en/evaluate/rfp-fact-sheet.md)
- [Questions to ask an AI support vendor, with Fini's answers](/en/evaluate/questions-to-ask.md)
- [Changelog](/en/changelog.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.