Skip to main content
An escalation policy is a short, written list of the conversations your AI agent must hand to a person, the signal that triggers each handoff, and what the person receives when it lands. Build it before launch from eight always-escalate categories (safety and crisis, identity and account takeover, fraud, legal and regulatory complaints, bereavement and vulnerable customers, high-value or irreversible actions, repeated failure, and an explicit request for a human), express each one as a trigger you can test, and then review the reasons behind every escalation on a fixed schedule. The goal is not the lowest escalation rate. It is that the agent resolves what it safely can, and a person gets the cases where a wrong answer would hurt the customer or the business.

Why write it down

A policy that only exists in people’s heads can’t be tested, drifts from the agent’s configuration, and leaves no record of what should have happened when a sensitive conversation goes wrong. A written policy gives support, compliance and engineering one document to agree on, and becomes the source of your test cases.

The always-escalate categories

Adapt the examples to your business, but think hard before removing a category.

A decision framework for everything else

For every other intent, ask how much harm a wrong outcome causes, and whether the agent has what it needs to get it right (approved knowledge, a working action, a clear policy). In the middle path, the agent gathers the details and drafts the reply or action, and a person checks it before it reaches the customer or your systems. It is often the right first step for a sensitive intent you plan to automate later.

How to express triggers

A trigger must be observable in the conversation or in your data. “Escalate angry customers” can’t be tested; “escalate when the customer mentions a regulator, ombudsman or lawyer” can. Put thresholds in data, not prose: a refund limit belongs in a rule that reads the order amount, not in a sentence the model interprets. When a topic has a defined workflow (cancel, refund, address change), let the workflow decide when to hand off instead of listing the whole topic as a trigger.

What to hand over

The person who picks up an escalation should never have to ask the customer to repeat themselves. Tell the customer what happens next, promise only what your team can meet, and stop the agent replying once a person is assigned.

Response-time expectations

Escalated cases are the hard ones, so give them their own targets per category rather than the queue default: a safety or account takeover case needs a much faster response than a refund over the limit. Decide what happens outside staffed hours for each category (for example an emergency resource for crisis cases, or an account freeze path for takeover reports), and make the handoff message match the target you can actually hit.

Testing the policy

Every category gets test conversations, and all of them run before every change to prompts, rules or knowledge.
  1. Write positive cases for each category, including indirect phrasing (“my mum passed away and her card is still being charged”).
  2. Write near misses that should not escalate, such as a customer asking how to report fraud in general, so you catch over-escalation too.
  3. Test multi-turn conversations. Many triggers only appear in the third or fourth message, after the agent has already started a workflow.
  4. Test in every language you support, and with typos, slang and mixed languages.
  5. Check the handover, not only the decision: the ticket lands in the right queue with the transcript, reason and attempted actions.

Reviewing escalation reasons over time

Tag every escalation with a reason from a fixed list, then read the distribution weekly. A useful taxonomy separates causes you fix in different places: Watch for two failure modes. Over-escalation shows up as policy escalations on topics the agent could handle, or customers asking for a person after the agent tried; read those conversations first. Under-escalation is invisible in escalation reports, because the conversation never escalated. Sample resolved conversations in your sensitive categories every week and check that none should have reached a person. Add every miss as a test case.

Policy template

Copy this table into your policy document and fill one row per trigger.

Common mistakes

  • Vague triggers that can’t be tested (“escalate frustrated customers”).
  • Escalating whole topics that have a working, rule-backed workflow.
  • No near-miss tests, so over-escalation is only found after launch.
  • Handoffs without context, so customers repeat themselves.
  • Only reading escalations, never the resolved conversations where one was missed.
For examples of ticket types that should go to a person, see the blog post When AI should escalate: 5 tickets to avoid.

Doing this in Fini

In Fini (usefini.com), each part of the policy maps to a specific setting:
  • Always-escalate categories: write them as plain-language triggers in the Escalation Topics subsection of the Planning Prompt. The defaults cover legal or regulatory issues and persistent human requests (agent_request_count >= 3, issue_repeat_count >= 3); add your business-specific categories on top. See Prompts.
  • Value thresholds and workflows: put limits in an Intent Rule branch that hands off by design, such as a refund above your limit. The branch ends in a handoff Reply that tells the customer a teammate will follow up. See Rulebook.
  • Topics that should always reach a person from the knowledge base: set the article-level Escalation field to Yes on those articles, so hitting them triggers an escalation path. See Articles.
  • Widget handoff into your helpdesk: a Business Rule with the On Escalation trigger creates the ticket in your helpdesk.
  • “Agent prepares, person approves” and staying silent: use an Internal Comment rule for sensitive intents, and No Reply with Human Agent Assigned Equals True so the agent never talks over a person. See Reply Rules.
  • Reviewing reasons: every escalation is tagged with a reason in four families (Knowledge, Action, User, Policy). Read the escalation doughnut and the Escalation reason filter in Analytics.
  • Handover and testing: how the handoff lands on each surface with the transcript, customer attributes and AI Steps trace is covered in Escalation and handoff. To test the policy, create Test Suite cases from real conversations for each category, and link them to a criteria group whose exact handoff check is marked Required to pass. Put near misses in a separate group that checks the agent did not hand off. Every fixed miss becomes a new case.

Escalation and handoff

How Fini decides to escalate, the reason taxonomy and the handoff on each surface.

Prompts

Escalation Topics in the Planning Prompt.

Reply Rules

No Reply, Internal Comment and Direct Reply conditions.

Preparing your APIs for an AI agent

Read and write endpoints, idempotency and confirmation before irreversible actions.