Concepts & Mental Model¶
This page is the mental model for guardrails — why they exist, where they fire in the agent run loop, how the executor runs them (blocking vs non-blocking), and how the implementation types relate to the safety concerns they cover. Read it first, then use the Guardrails reference, Responsible AI, and Managed governance for depth.
Why guardrails exist¶
An agent takes untrusted input, calls a model, and acts on the world through tools. Each of those boundaries is a place something can go wrong: a prompt injection in the input, PII or secrets in the output, an unsafe argument to a tool. A guardrail is an assertion placed at one of those boundaries — it inspects the data and either lets it pass or blocks the run.
Guardrails assert; middleware transforms
A guardrail is pass/fail — it validates and can raise. Middleware changes the data flowing through (trim history, redact, rewrite). Use a guardrail for a policy check that should block on failure; use middleware when you want to modify what flows through the loop.
The four positions¶
Guardrails attach at four positions — the boundaries the agent run loop crosses:
| Position | Fires | Guards against |
|---|---|---|
input |
Before the model sees the user input | Prompt injection, off-topic/abuse, disallowed requests |
tool_call |
On a tool's arguments, before it runs | Unsafe/destructive tool arguments |
tool_result |
On a tool's output, after it runs | Leaking sensitive data a tool returned |
output |
On the final answer, after the loop ends | PII, secrets, toxicity, hallucination, off-policy replies |
Think of a run as data flowing through up to four gates: input → (loop:
tool_call → tool_result, per tool call) → output. Positions are
independent — use any combination. You attach them all the same way:
Agent(guardrails=[...]), and each guardrail declares its own position.
The execution model¶
At each position the executor (fastaiagent/guardrail/executor.py) runs the
applicable guardrails with a deliberate two-phase strategy:
- Blocking guardrails run first, sequentially. The first one that fails
raises
GuardrailBlockedErrorimmediately — the run stops and nothing after it executes (fail-fast). - Non-blocking guardrails run after, in parallel (
asyncio.gather). A non-blocking failure is recorded but does not stop the run, and an exception that escapes one is caught and turned intoGuardrailResult(passed=False, errored=True, action_taken="blocked")rather than crashing the agent.
That catch is not fail-open
The recorded verdict is a failure, not a pass — the executor fails
closed on what it writes down, and on_error gets no say on this path
(executor.py builds the result directly). What it does not do is halt
the run, because a non-blocking rule never can. "Does not stop the run" and
"passes the payload" are different things; only the first is true here.
Verified against a live run
With one blocking + two non-blocking guardrails at input: on clean data
the order was blocking → then both non-blocking, and a non-blocking
"fail" was recorded without stopping the run. When the blocking guardrail
failed, it raised GuardrailBlockedError and the non-blocking guardrails
never ran — fail-fast, as designed.
So the mental model is: blocking = a gate that can stop the run; non-blocking
= an observer that records but never blocks. Set blocking=True for policy
you must enforce, blocking=False for signals you want to watch.
Concretely, execute_guardrails(guardrails, data, position) filters the list to
that position, then:
for g in blocking: # sequential, fail-fast
result = await g.aexecute(outcome.data)
if halts(g, result): # ← what the action actually did
raise GuardrailBlockedError(...) # stops the run
results += await asyncio.gather( # non-blocking, parallel
*[g.aexecute(outcome.data) for g in non_blocking],
return_exceptions=True, # → GuardrailResult(passed=False, errored=True)
)
A blocking failure at any position raises GuardrailBlockedError, which
propagates out of the agent run — that's how input/tool_call/tool_result/output
all "stop" the run when they must.
The third axis: what a failure costs¶
blocking decides whether a rule runs inline and can halt. It does not decide
what a failure is worth. That is action, a third and independent axis:
| Axis | Field | Question |
|---|---|---|
| Scheduling | blocking |
Does this run inline, and can it halt? |
| Degradation | on_error |
What does an un-runnable check mean? |
| Consequence | action |
What does a genuine failure cost? |
action is one of block (the default), warn, mask, override or reask.
A warn rule still runs inline and the caller still waits for its verdict — it
just doesn't stop; a mask rule redacts the offending span and the run carries
on with the redacted value. Because mask and override change the payload,
execute_guardrails returns a GuardrailOutcome carrying both the verdicts and
the value to carry forward.
The rule that keeps this honest: every caller branches on what the action
actually did (action_taken), never on what it was configured to do. That
is why "an errored check always blocks" and "a mask with nothing to mask blocks"
need no special case anywhere.
See Actions, severity & floor for the full contract, including where a rewrite cannot be applied and the run blocks instead.
The verdict object¶
Every guardrail — whatever its type — resolves to one
GuardrailResult(passed, score, message, execution_time_ms, metadata, errored,
action, action_taken, modified_data). The executor branches on passed and
action_taken; score, message and metadata are for observability (the
Local UI reads them), and modified_data carries the rewritten payload when the
action produced one. For a code guardrail, your fn can
return a bare bool (coerced to GuardrailResult(passed=...)) or a full
GuardrailResult. A guardrail crash never crashes the agent — a raised
exception is caught and resolved according to the guardrail's on_error policy
(see below), with errored=True so a degraded result is never mistaken for a
real verdict.
When the check itself fails: on_error¶
A model-judged guardrail (toxicity_check(mode="llm"), grounded, an
llm_judge, OpenAI moderation, …) depends on an LLM call that can fail —
a timeout, a 5xx, an unparseable response. That is different from the check
running and returning a verdict, and you get to decide what it means:
toxicity_check(mode="llm", on_error="block") # fail closed: an error blocks
toxicity_check(mode="llm", on_error="allow") # fail open: an error passes through
on_error="block"(fail closed) — an errored check is treated as a failure. Use it for policy you must enforce even when the checker is down (a bank would rather block than guess). This is the default for the guardrails that already behaved this way (grounded,openai_moderation,allowed_topics, and anyGuardrail(...)/llm_judgeyou build yourself).on_error="allow"(fail open) — an errored check lets the content through. Use it when availability beats strictness (a high-traffic chatbot would rather serve than let a flaky moderation API take it offline). This is the default for the convenience classifiers that already behaved this way (toxicity_check,no_prompt_injection,banned_topics).
Either way the outcome is visible: the result carries errored=True, the
trace span records a guardrail.errored attribute and an "error" check
result, and the Local UI logs the event with an errored outcome — so you can
see exactly how often a guardrail is degrading instead of guarding. Guardrail
LLM calls also get a small automatic retry, so a single transient blip doesn't
trip the policy at all.
on_error answers "could not run", not "could not apply"
Applying the action can fail too — a mask re-runs a detector or a regex
substitution to build the redacted payload, and that work can raise. Since
1.62.0 that is caught, but it is caught separately and on_error gets no
say: the check did run and returned a verdict, and only the consequence
could not be applied, so the outcome is a block even under
on_error="allow". It is the same rule as a mask that finds no span to
redact — a mask that raised is strictly worse than one that found nothing.
The result carries errored=True, action_taken="blocked", and an
action_error in its metadata.
Before 1.62.0 this was not caught at all: apply_action sat outside
run_guardrail's handler, so a failure there surfaced as a bare
AttributeError out of agent.run() — no errored flag, no on_error, not
even a GuardrailBlockedError.
The no_pii(), no_secrets(), json_valid() and allowed_domains() builtins
don't make a fallible call, so on_error doesn't come up for them — they are
reliable hard blocks. The pii type is different: it raises on an unknown
entity, an unknown backend, or backend="presidio" without the [safety]
extra, and on_error decides what each costs. Detection that cannot run is not
detection that found nothing.
How each type decides¶
run_guardrail dispatches on GuardrailType to ten deciders, all producing the
same GuardrailResult:
| Type | How it decides |
|---|---|
code |
Runs your Python fn(data) — arbitrary logic |
regex |
Searches for config["pattern"]; should_match flips whether a match means pass (True) or fail (default False) |
schema |
Validates the data against a JSON Schema |
llm_judge |
Calls a model with a rubric and parses the verdict fail-closed (ambiguous → fail) |
classifier |
Keyword matching, not a model: substring-searches config["categories"] (a {category: [keyword, …]} map) and blocks the hits, narrowed by config["blocked"] when one is given |
content_safety |
Scores the payload against the MLCommons hazard taxonomy, with a bar per category |
groundedness |
Scores an answer against the context it was given |
topic |
Classifies against named topics, then deny (blocklist) or allow (on-topic gate) |
pii |
Detects personal data with the shared detectors; can redact, not just refuse |
secrets |
Detects leaked credentials and tokens; can redact, not just refuse |
Three of them — content_safety, groundedness and topic — are model-backed
judges with structure; two more — pii and secrets — are detector-backed and
are the only types that can genuinely redact. See
Actions, severity & floor for the
first three and its detector section
for the other two.
classifier is not a model, and since 1.64.0 blocked only narrows
classifier is pure substring matching over the keyword lists in
config["categories"] — there is no endpoint, no model and no score
threshold. Two consequences worth knowing:
- A rule with no categories raises rather than reporting a pass: a rule that scans for nothing cannot tell a clean payload from a dirty one.
- Until 1.64.0, a rule with no
blockedlist detected categories and then reported success — nothing ever blocked. It now blocks every category it detects, matching the plane;blockednarrows that set, and its absence no longer disarms the rule.
This is the mechanical basis for the two-axis view below: the type is which decider runs; the concern is what you point it at.
Two axes: implementation type × what it checks¶
A guardrail is described by two independent things — don't conflate them:
- Implementation type (
GuardrailType) — how it decides:code,regex,schema,llm_judge,classifier,content_safety,groundedness,topic,pii,secrets. - What it checks — the concern: prompt injection, PII, secrets, toxicity,
groundedness, topic, moderation. The Responsible AI
bundle is a curated set of these, each implemented as an ordinary
Guardrail.
The same concern can be implemented different ways, with a real cost trade-off:
code/regex/schema are free and instant; llm_judge/classifier and
LLM-backed safety checks cost an inference call but catch things patterns
can't. Reach for cheap deterministic checks first, LLM-backed ones where
nuance matters.
Composition¶
- Stack them — put several guardrails on one agent; the executor groups them by position and applies the blocking/non-blocking rules per position.
- Bundle them —
responsible_ai(...)returns a list ofGuardrails you spread intoguardrails=[...]; see Responsible AI. - Govern them centrally — a connected agent can defer high-stakes tool calls to a platform policy that can require human approval; see Managed governance. That path pauses the run rather than simply passing/failing.
What a guardrail run puts on the trace¶
Every guardrail that runs emits one child span — on a pass as well as a
block, so the console shows green checks and not only failures. The span is
classified with the OpenInference standard kind and carries the outcome in the
fastaiagent.guardrail.* namespace:
| Attribute | Meaning |
|---|---|
openinference.span.kind |
Always "GUARDRAIL" — the standard classifier |
fastaiagent.guardrail.name |
The guardrail's name. How the platform resolves the span to a guardrail |
fastaiagent.guardrail.position |
input / tool_call / tool_result / output |
fastaiagent.guardrail.passed |
The verdict |
fastaiagent.guardrail.errored |
true when the check couldn't run and passed reflects on_error, not a real verdict |
fastaiagent.guardrail.checks |
JSON: [{"name": ..., "result": "pass"|"block"|"error"}] |
fastaiagent.guardrail.action |
What the rule was configured to cost: block / warn / mask / override / reask |
fastaiagent.guardrail.action_taken |
What it actually did: none / blocked / warned / masked / overridden / reask |
fastaiagent.guardrail.severity |
low / medium / high / critical, when the rule carries one |
fastaiagent.guardrail.floor |
true when the rule is the domain-wide baseline |
The split matters: OpenInference standardizes the span kind, not the outcome
fields. There is no ecosystem convention for "what did this guardrail decide",
so fastaiagent.guardrail.* is ours, and it rides under the standard kind.
Anything that understands OpenInference recognizes the span as a guardrail;
anything that understands FastAIAgent additionally reads the verdict.
The span status follows the verdict — OK on pass, ERROR on block — with one
deliberate exception: a degraded pass (errored=true with on_error="allow")
keeps an OK status, because the run did continue. The errored attribute is
what tells a fail-open apart from a genuine pass.
Legacy span_type marker
Guardrail spans also carry span_type="guardrail" alongside the standard
kind. That dual-write exists for platform deployments predating the
OpenInference reader and is transitional — don't build on it.
If you're enforcing guardrails from a runtime that isn't fa.Agent, you emit
this same span yourself with fa.emit_guardrail(...). See Guardrails and evals
without the runtime.
Next steps¶
- Guardrails — the full reference: all ten types, built-in factories, custom guardrails, serialization
- Actions, severity & floor — what a failure costs, plus the three model-backed and two detector-backed check types
- Guardrails & evals without the runtime — borrowing
run_guardrailfrom a foreign framework - Responsible AI (Trust Layer) — the safety bundle by concern
- Managed governance — platform-enforced, approval-gated tool policy
- Agents — the run loop — exactly where each position fires