Skip to main content
Guardrails check generated replies before delivery. Open Guardrails from the admin sidebar and select an agent. These policies are separate from prompt guidelines, QA tags, and Reply Behavior.

Add and configure a guardrail

1

Choose a check

Click Add guardrail and select a check type. Each built-in type can be added once per agent; you can add up to 10 custom rules.
2

Set its definition

Enter the terms, attribute keys, domains, or evaluation criteria for the check. Click Add guardrail to save a new definition, or Save changes when editing one. Checks with no configurable fields are added directly.
3

Set scope and enablement

Use Scope and Add channel to select channels. No selected channels means All channels. The ON / OFF switch enables or disables the saved policy. Scope and enablement changes save immediately.
New policies are enabled by default. Configuration changes affect the saved policy directly; there is no separate publish step.

Check types

Reply outcomes

A passing reply is sent unchanged. When a check finds a violation, Fini attempts one rewrite and checks the result again. If the rewrite remains unsafe or cannot preserve the original reply, Fini replaces the regular reply with a handoff message and marks the conversation for escalation. A blocked inactivity follow-up is skipped instead of sending a handoff message.
Guardrails are not a fail-closed security boundary. Check errors or timeouts can leave the original reply unchanged. An individual check can report could not run even when the overall outcome is Passed. Review individual verdicts as well as the overall outcome.
Channel scope does not exclude dashboard tests, Test Suite, or replay evaluation: those sources remain eligible regardless of the selected production channels. Standalone widget conversations use the widget scope.

Review activity

Use All, Firing, Quiet, or Off to filter the policy list. Activity windows are 24h, 7d, and 30d. Hits count initial failed checks, including replies that were successfully rewritten; they are not a count of escalations. Open a policy’s activity in Inbox, or use the Guardrail filter under Quality. In AI Steps, inspect the Guardrails section for policy verdicts and the reply outcome: Passed, Rewritten, Escalated to a human, or Check failed, sent as generated.

Why a guardrail is not behaving as expected

Check that it is enabled, applies to the conversation channel, and has the intended definition. A hit means an initial failed check, not every evaluation.
Dashboard tests, Test Suite, and replay are exempt from channel filtering. This lets you evaluate policies before using them on a production channel.
Inspect the per-policy verdicts in AI Steps. Guardrail evaluation and rewrite errors can send the original reply; do not treat these outcomes as proof that every check passed.
For programmatic configuration and reporting, see the Guardrails API.