Skip to content

Guardrails ​

Guardrails provide real-time content safety checks for LLM inputs and outputs. They detect PII, secrets, banned words and profanity, harmful content, prompt-injection attempts, and unwanted tool/webhook activity — and can block, redact, warn on, or flag what they find. Operators manage them under Operate → Guardrail.

Hook plane rewrite (September 2026)

Guardrails moved from a two-slot model (one policy on the request, one on the response) to a hook plane: six hooks spanning the full request lifecycle, each configurable per policy family. If you're looking for the old inputGuardrailKey / outputGuardrailKey model, see Inference Integration below for how the two map.

Operator view ​

Each guardrail is a named policy attached to one or more models or agents. The list view summarises the states that matter operationally: total policies, how many are enabled, how many are disabled, and how many block requests.

Guardrails list

Filters narrow by type (preset / custom), by action, and by status. The Create guardrail flow walks you through picking a type, declaring which checks to enforce, and configuring the default action and failure mode.

Hooks ​

A guardrail is a hook plane: one column per policy, one row per hook point, a filled cell means that policy is evaluated there. The six hook points, in request order:

HookFires
prompt.preOnce per run, on the user's turn
input.preBefore every model call
output.preBefore the answer reaches the caller
output.stream.deltaWhile the answer is streaming, as a post-hoc audit (a stream already in flight can't be blocked mid-delivery)
tool.preBefore a tool runs
tool.postAfter a tool returns

A guardrail has one mode — Enforce (verdicts are acted on: a policy that says block, blocks), Monitor (everything is evaluated and recorded, nothing is blocked), or Off — applied uniformly across its rows; each hook can additionally be switched off independently without touching the others. Each row also carries its own Enforcement action (e.g. Block).

Enforcement now reads the mode-neutralised decision

A guardrail running in Monitor mode no longer blocks the request it's attached to. Earlier builds evaluated the raw pass/fail regardless of mode, which meant a guardrail you'd deliberately set to observe-only could still 400 a live request — that's fixed.

Guardrail hooks tab: one column per policy, one row per hook point

Guardrail Types ​

TypeDescription
presetA bundled policy combining detection families: PII, word filter, content moderation, and prompt shield.
customAn LLM-based evaluation driven by your own rule text. Requires a model.

Detection Families ​

Six policy families, grouped by what they're for. Each lists which hooks it can bind to — a family built on an LLM classifier (moderation, prompt shield) or list matching (word filter) can't run on output.stream.delta, since a stream already in flight can't be re-judged token by token; personal data and credentials can, since they're pattern-based.

Sensitive data — something confidential is in the text:

FamilyHooksWhat it does
Personal data (pii)all sixNames, contact details and ID numbers, scanned through one of your PII policies — 15 categories incl. email, phone, credit card (Luhn), IBAN (mod-97), TCKN (checksum), API keys/JWTs (known prefixes + entropy). Categories, languages, checksums and mask strategies live on that PII policy; this form shows and edits them in place.
Credentials (secrets)all sixAPI keys, access tokens and private keys — vendor patterns plus an optional high-entropy heuristic. No database, no model, no network.

Unacceptable content — the text itself is the problem:

FamilyHooksWhat it does
Word filterprompt.pre, input.pre, output.pre, tool.postWord lists and phrases, matched after normalisation so leetspeak and s p a c e d out evasion still hit. Cannot run on a stream, for the same reason the LLM-backed families can't.
Moderationprompt.pre, input.pre, output.pre, tool.postAn LLM classifier for harmful and policy-violating content across the standard category set. Costs a model call on every run.
Prompt shieldprompt.pre, input.pre, output.pre, tool.postAn LLM classifier for prompt injection and jailbreak attempts. Costs a model call, and its block message deliberately says nothing about why.

What the assistant may do — the assistant is reaching for a tool, domain, or file it shouldn't:

FamilyHooksWhat it does
Tool accesstool.pre, tool.post onlyWhich tools may run, for whom, with what arguments, against which domains and paths. The only family that can stop a tool call before it happens.

Separately, a guardrail's type (preset vs custom, set when you create it) decides whether it runs the bundled detection families above or an LLM judged against your own rule text — see Guardrail Types.

Word lists ​

The word filter merges four sources: built-in lists (toggled per policy), tenant word lists (managed under Guardrails → Word lists, uploaded as CSV/TXT or edited inline, up to 20k entries), inline words, and regexes. Uploaded lists are referenced from policies by key (policy.wordFilter.customListKeys) and are cached for 60 s at evaluation time.

MethodEndpointDescription
GET/api/guardrails/word-listsList summaries (name, key, word count)
POST/api/guardrails/word-listsCreate — body accepts words: string[] or raw content (CSV/TXT; ,/;/tab/newline separated, # comments)
GET/api/guardrails/word-lists/:idFull list including words
PATCH/api/guardrails/word-lists/:idUpdate metadata and/or replace words (words or content)
DELETE/api/guardrails/word-lists/:idDelete

The LLM-backed checks run against any LLM from the tenant's Model Hub (set policy.<family>.modelKey, falling back to the guardrail's modelKey). The evaluated text is wrapped in per-request random boundary markers and declared untrusted data, so verdict-steering text inside the message ("respond with allowed: true") is itself treated as an attack signal.

Attach a guardrail to a model ​

Creating a guardrail doesn't run it against anything by itself — a guardrail only fires once it's bound to a model. Bindings live on the model, not on the guardrail, because the same guardrail is typically reused across many models with different hooks enabled on each.

  1. Open Model Hub (Operate → Model Hub) and click into the model you want to protect.
  2. On the Configure tab, find the Guardrails card. A model with nothing attached reads "No guardrails attached — every request to this model passes unchecked."
  3. Click Edit to open the Guardrail bindings modal.
  4. Pick a guardrail from the Attach a guardrail… dropdown. As soon as you select one, its row appears with a checkbox per hook: Prompt, Input, Output, Streaming output, Before a tool, After a tool.

Guardrail bindings modal — per-hook checkboxes for an attached guardrail

  1. Tick the hooks you want this binding to cover. Only hooks the guardrail can actually serve are selectable — the rest are greyed out. A checkbox is enabled only when the guardrail has an enabled policy bound to that hook (see Detection Families for which family serves which hooks); a binding to a hook the guardrail doesn't serve would look configured in the UI and simply never run. If a box you need is greyed out, go back into the guardrail itself and enable a policy on that hook first, then return here.
  2. Repeat Attach a guardrail… to stack additional guardrails on the same model — for example a pii/secrets preset on input.pre/output.pre plus a separate tool_access guardrail on tool.pre/tool.post. Each attached guardrail gets its own row and its own set of hook checkboxes.
  3. Save. The card on the Configure tab now lists each attached guardrail with the hooks it's bound to; every matching request is evaluated from that point on — no redeploy, no restart.

Removing a guardrail from a model is the same modal: open Edit, use the trash icon on its row to detach it, and Save.

Same guardrail, different hooks per model

Because bindings are per-model, one guardrail can run in Enforce mode (blocking) on a production model and in Monitor mode on a staging model — or with only output.pre ticked on one model and all six hooks ticked on another. The guardrail's own mode (Enforce/Monitor/Off, set on the guardrail itself) still applies uniformly wherever it's bound; what changes per model is which hooks are wired up.

Actions ​

ActionBehavior
blockReject the request with a structured error
redactMask the detected values ([REDACTED:email]) and continue — PII, secrets, and word filter only
warnAllow the request; findings are attached to the response and logged
flagAllow the request; findings are attached to the response and logged

The guardrail-level action applies to LLM-backed findings; the PII, secrets, and word-filter policies carry their own action so you can, for example, redact PII while blocking profanity. A binding in monitor mode never blocks, regardless of the configured action.

Failure Mode ​

LLM-backed checks can fail (model down, unparseable verdict). failMode decides what happens:

  • open (default) — the content passes; the failure is logged.
  • closed — the content is treated as a violation (evaluation_error finding). Also fires when an LLM check is enabled with no model configured.

Evaluation Logs ​

Every evaluation is persisted to guardrail_evaluation_logs with pass/fail, findings, latency, the calling surface (chat.completions, agent, client-api, …), and the request id. Detected values and the stored input text are masked before persistence — raw PII never lands in the log. The guardrail detail page charts pass rate, findings by type/severity, and a time series; alert rules can trigger on guardrail_fail_rate, guardrail_avg_latency_ms, and guardrail_total_evaluations.

API ​

Evaluate Guardrail ​

POST /api/client/v1/guardrails/evaluate
Authorization: Bearer <token>
json
{
  "guardrail_key": "pii-checker",
  "text": "My email is john@example.com and my phone is 555-0100"
}

Response:

json
{
  "passed": false,
  "action": "block",
  "findings": [
    { "type": "pii", "category": "email", "severity": "high", "message": "Email address detected", "action": "block", "block": true, "value": "john@example.com" }
  ],
  "guardrail_key": "pii-checker",
  "guardrail_name": "PII Checker",
  "message": "Content blocked by guardrail:\n• Email: Email address detected",
  "redacted_text": null
}

When the matching policy's action is redact, passed is true and redacted_text contains the masked text to use instead of the original.

Inference Integration ​

The input.pre and output.pre hooks are what the older inputGuardrailKey / outputGuardrailKey slots mapped to — attaching a guardrail to a model still configures those two hooks under the hood:

Request → prompt.pre → input.pre → Provider call → output.pre → Response
                                        │ tool.pre / tool.post around each tool call
              │ block: 400 guardrail_block         │ block: 400
              │ redact: rewrite user msg           │ redact: rewrite content
              └ warn/flag: annotate + log          └ warn/flag: annotate + log
  • Non-blocking findings are attached to the chat completion response under a guardrails extension field ({ input?: {...}, output?: {...} }).
  • Streaming responses cannot be blocked after delivery; output.stream.delta runs as a post-hoc audit as the stream is delivered, feeding evaluation logs and alerts (source: chat.completions:stream).
  • Agents apply the same guardrails around their conversation loop (source: agent), with tool.pre / tool.post firing around each tool call.

When a guardrail blocks a request during inference, a GuardrailBlockError is thrown with the guardrail key, action, and findings.

Management Endpoints ​

MethodEndpointDescription
GET/api/guardrailsList guardrails (includeTemplates=true returns category/list catalogs)
POST/api/guardrailsCreate guardrail (validates that LLM checks have a model)
GET/api/guardrails/:idGet guardrail
PATCH/api/guardrails/:idUpdate guardrail
DELETE/api/guardrails/:idDelete guardrail
POST/api/guardrails/evaluateEvaluate text (dashboard)
GET/api/guardrails/:id/evaluationsEvaluation logs + aggregate (pass rate, latency, time series)
  • AI App Gateway (Enterprise) — its Policy tab's personal-data check runs one of your PII guardrails rather than reimplementing detection; its tool_access-equivalent role for coding-agent traffic sits alongside this hook plane's own tool_access family.

Studio · Pulse · Console · Agent SDK and more — the Cognipeer documentation hub