> ## Documentation Index
> Fetch the complete documentation index at: https://docs.usefini.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Guardrails

> Configure per-agent reply checks, channel scope, and review guardrail activity.

Guardrails check generated replies before delivery. Open **Guardrails** from the admin sidebar and select an agent. These policies are separate from [prompt guidelines](/en/configuration/prompts), [QA tags](/en/configuration/tags), and [Reply Behavior](/en/automations/reply-behavior).

## Add and configure a guardrail

<Steps>
  <Step title="Choose a check">
    Click **Add guardrail** and select a check type. Each built-in type can be added once per agent; you can add up to 10 custom rules.
  </Step>

  <Step title="Set its definition">
    Enter the terms, attribute keys, domains, or evaluation criteria for the check. Click **Add guardrail** to save a new definition, or **Save changes** when editing one. Checks with no configurable fields are added directly.
  </Step>

  <Step title="Set scope and enablement">
    Use **Scope** and **Add channel** to select channels. No selected channels means **All channels**. The **ON / OFF** switch enables or disables the saved policy. Scope and enablement changes save immediately.
  </Step>
</Steps>

New policies are enabled by default. Configuration changes affect the saved policy directly; there is no separate publish step.

## Check types

| Check | Configuration and behavior |
| - | - |
| Internal reasoning leak | Checks for internal reasoning, prompt text, and pipeline terminology using a built-in term list and an LLM check. Add extra terms or patterns if needed. |
| Banned terms | Case-insensitive whole-word matching for terms and phrases, plus optional regular expressions. Supply at least one term or pattern. |
| Confidential attributes | Select user attribute keys whose values should not appear in replies. This is a value-matching check, not blanket detection of all personal data. |
| URL allowlist | Supply bare domains, without a scheme or path. Subdomains of allowed domains are allowed too. |
| AI disclosure | Checks for replies describing the assistant as an AI or language model, or naming an AI provider or model. No definition fields are required. |
| Custom rule | Give the rule a name and evaluation instruction. Optional pass/fail examples help calibrate the LLM check. |

## Reply outcomes

A passing reply is sent unchanged. When a check finds a violation, Fini attempts one rewrite and checks the result again. If the rewrite remains unsafe or cannot preserve the original reply, Fini replaces the regular reply with a handoff message and marks the conversation for escalation. A blocked inactivity follow-up is skipped instead of sending a handoff message.

<Warning>
  Guardrails are not a fail-closed security boundary. Check errors or timeouts can leave the original reply unchanged. An individual check can report **could not run** even when the overall outcome is **Passed**. Review individual verdicts as well as the overall outcome.
</Warning>

Channel scope does not exclude dashboard tests, Test Suite, or replay evaluation: those sources remain eligible regardless of the selected production channels. Standalone widget conversations use the widget scope.

## Review activity

Use **All**, **Firing**, **Quiet**, or **Off** to filter the policy list. Activity windows are **24h**, **7d**, and **30d**. Hits count initial failed checks, including replies that were successfully rewritten; they are not a count of escalations.

Open a policy's activity in [Inbox](/en/testing/inbox), or use the **Guardrail** filter under **Quality**. In **AI Steps**, inspect the **Guardrails** section for policy verdicts and the reply outcome: **Passed**, **Rewritten**, **Escalated to a human**, or **Check failed, sent as generated**.

## Why a guardrail is not behaving as expected

<AccordionGroup>
  <Accordion title="The policy has no hits">
    Check that it is enabled, applies to the conversation channel, and has the intended definition. A hit means an initial failed check, not every evaluation.
  </Accordion>

  <Accordion title="A test still runs a channel-scoped policy">
    Dashboard tests, Test Suite, and replay are exempt from channel filtering. This lets you evaluate policies before using them on a production channel.
  </Accordion>

  <Accordion title="A reply was sent despite a check error">
    Inspect the per-policy verdicts in AI Steps. Guardrail evaluation and rewrite errors can send the original reply; do not treat these outcomes as proof that every check passed.
  </Accordion>
</AccordionGroup>

For programmatic configuration and reporting, see the [Guardrails API](/en/api-reference/guardrails).


## Related topics

- [Update guardrail policy](/en/api-reference/update-guardrail-policy.md)
- [Get guardrail hits](/en/api-reference/get-guardrail-hits.md)
- [Get guardrail policy](/en/api-reference/get-guardrail-policy.md)


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.