Agents¶
An Agent is the core building block of the FastAIAgent SDK. It wraps an LLM with tools, guardrails, and memory to create an autonomous assistant that can reason, take actions, and validate its own output.
Creating an Agent¶
from fastaiagent import Agent, LLMClient
agent = Agent(
name="support-bot",
system_prompt="You are a helpful customer support agent.",
llm=LLMClient(provider="openai", model="gpt-4.1"),
)
result = agent.run("How do I reset my password?")
print(result.output)
Supported LLM providers:
| Provider | Example |
|---|---|
| OpenAI | LLMClient(provider="openai", model="gpt-4.1") |
| Anthropic | LLMClient(provider="anthropic", model="claude-sonnet-4-6") |
| Ollama | LLMClient(provider="ollama", model="llama3") |
| Azure | LLMClient(provider="azure", model="gpt-4", base_url="https://myendpoint.openai.azure.com/openai/deployments/gpt-4/") |
| AWS Bedrock | LLMClient(provider="bedrock", model="anthropic.claude-3-sonnet-20240229-v1:0") |
| Custom | LLMClient(provider="custom", model="my-model", base_url="https://my-api.com/v1") |
Agent with Tools¶
Tools let your agent take actions — call APIs, search databases, run calculations.
from fastaiagent import Agent, FunctionTool, LLMClient
def get_weather(city: str) -> str:
"""Get current weather for a city."""
return f"Sunny, 22°C in {city}"
def search_orders(order_id: str) -> str:
"""Look up an order by ID."""
return f"Order {order_id}: Shipped, arriving tomorrow."
agent = Agent(
name="assistant",
system_prompt="You help customers with weather and order questions. Use tools.",
llm=LLMClient(provider="openai", model="gpt-4.1"),
tools=[
FunctionTool(name="get_weather", fn=get_weather),
FunctionTool(name="search_orders", fn=search_orders),
],
)
result = agent.run("What's the weather in Paris and where is order ORD-123?")
print(result.output) # LLM's final text response
print(result.tool_calls) # List of tool calls made
print(result.tokens_used) # Tokens across EVERY turn of the run, not just the last
print(result.latency_ms) # Total execution time
How tool calling works:
1. Agent sends messages + tool schemas to the LLM
2. LLM decides to call one or more tools (or respond directly)
3. SDK executes the tools and sends results back to the LLM
4. LLM generates a final response using the tool results
5. This loop repeats up to max_iterations times
The @tool Decorator¶
For quick tool creation:
from fastaiagent.tool import tool
@tool(name="calculate")
def calculate(expression: str) -> str:
"""Evaluate a math expression."""
return str(eval(expression))
# Use directly — it's a FunctionTool
result = calculate.execute({"expression": "2 + 2"})
Tool Types¶
| Type | Use Case | Example |
|---|---|---|
FunctionTool |
Wrap any Python function | FunctionTool(name="calc", fn=my_func) |
RESTTool |
Call an HTTP API | RESTTool(name="weather", url="https://api.weather.com/v1", method="GET") |
MCPTool |
Connect to MCP server | MCPTool(name="search", server_url="http://localhost:3000") |
Agent with Guardrails¶
Guardrails validate input/output at every step, blocking unsafe content automatically.
from fastaiagent import Agent, LLMClient
from fastaiagent.guardrail import no_pii, toxicity_check, json_valid
agent = Agent(
name="safe-bot",
system_prompt="You are a helpful assistant.",
llm=LLMClient(provider="anthropic", model="claude-sonnet-4-6"),
guardrails=[
no_pii(), # Blocks SSN, email, phone, credit cards in output
toxicity_check(), # Blocks toxic language
],
)
result = agent.run("What are the benefits of eating healthy?")
print(result.output) # Clean output passes both guardrails
Important:
no_pii()is an output guardrail — it checks the LLM's response, not the user's input. If the LLM happens to include a real SSN, email, or phone number in its response, the guardrail blocks it and raisesGuardrailBlockedError. Most LLMs will refuse to output real PII on their own, so this guardrail acts as a safety net for edge cases, tool results that leak PII, or less-guarded models.
To guard against input containing PII, set the position explicitly:
from fastaiagent import Agent, LLMClient
from fastaiagent._internal.errors import GuardrailBlockedError
from fastaiagent.guardrail import no_pii, GuardrailPosition
agent = Agent(
name="input-safe-bot",
llm=LLMClient(provider="openai", model="gpt-4.1"),
guardrails=[
no_pii(position=GuardrailPosition.input), # Blocks PII in user input
no_pii(position=GuardrailPosition.output), # Blocks PII in LLM output
],
)
# User input containing an SSN is blocked before reaching the LLM
try:
result = agent.run("My SSN is 123-45-6789, can you store it?")
except GuardrailBlockedError as e:
print(f"Blocked: {e}") # "PII detected: SSN"
Built-in guardrail factories:
All thirteen are exported from fastaiagent.guardrail. The default position is
in the signature — no_pii() is an output guardrail, no_prompt_injection()
an input one, allowed_domains() a tool_call one — and every factory
takes position= to move it.
| Factory | What it checks | Default position |
|---|---|---|
no_pii(entities=…, backend=…) |
Email, US phone, SSN, credit cards (Luhn-validated) | output |
no_secrets() |
Leaked credentials, API keys and tokens | output |
no_prompt_injection(mode="heuristic"|"llm") |
Prompt-injection / jailbreak attempts. on_error="allow" by default |
input |
json_valid() |
Output is valid JSON | output |
toxicity_check(mode="keyword"|"llm") |
Toxic language. on_error="allow" by default |
output |
openai_moderation(model="omni-moderation-latest") |
Content flagged by the OpenAI moderation endpoint. on_error="block" |
output |
grounded(reference, threshold=0.7) |
Whether the answer is supported by reference. on_error="block" |
output |
no_hallucination(reference, threshold=0.7) |
Alias of grounded() — same check, the name you reach for |
output |
banned_topics([...]) |
Content falling under any banned topic (blocklist). on_error="allow" |
output |
allowed_topics([...]) |
Content outside the given topics (allowlist gate). on_error="block" |
output |
cost_limit(max_usd=0.10) |
The run's accumulated LLM spend so far. Blocks over budget; raises (so on_error decides) when the model has no rate in the pricing table, because an unpriced run is not a free one. Until 1.67.0 this always passed. |
output |
allowed_domains(["api.example.com"]) |
URL domains in tool calls | tool_call |
responsible_ai(...) |
Not one guardrail — returns a list to spread into guardrails=[...]. See Responsible AI |
(several) |
See the Guardrails reference for each factory's full
signature and on_error
for what a degraded check costs.
Custom guardrails:
from fastaiagent.guardrail import Guardrail, GuardrailPosition, GuardrailType
# Inline function
guardrail = Guardrail(
name="max_length",
position=GuardrailPosition.output,
blocking=True,
fn=lambda text: len(text) < 500,
)
# Regex-based
guardrail = Guardrail(
name="no_urls",
guardrail_type=GuardrailType.regex,
position=GuardrailPosition.output,
config={"pattern": r"https?://", "should_match": False},
)
Guardrail positions: input, output, tool_call, tool_result
Blocking modes:
- blocking=True — raises GuardrailBlockedError if validation fails
- blocking=False — logs the failure but continues execution
Agent Configuration¶
from fastaiagent import Agent, AgentConfig
agent = Agent(
name="configured-agent",
llm=LLMClient(provider="openai", model="gpt-4.1"),
config=AgentConfig(
max_iterations=5, # Max tool-calling loop iterations (default: 10)
tool_choice="auto", # "auto", "required", "none"
temperature=0.7, # LLM temperature override
max_tokens=1000, # Max response tokens
),
)
Streaming¶
Stream tokens from the agent as they are generated, rather than waiting for the full response:
from fastaiagent.llm.stream import TextDelta, ToolCallStart
async for event in agent.astream("What's the weather in Paris?"):
if isinstance(event, TextDelta):
print(event.text, end="", flush=True)
elif isinstance(event, ToolCallStart):
print(f"\n[Calling {event.tool_name}...]")
A sync wrapper is also available:
Streaming runs input guardrails before streaming begins and output guardrails after streaming completes. Memory is updated at the end.
See Streaming for full details, event types, and chat UI patterns.
Sync vs Async¶
Every method has both sync and async versions:
Streaming also has both forms:
# Async — yields events in real time
async for event in agent.astream("Hello"):
...
# Sync — collects into AgentResult
result = agent.stream("Hello")
The sync run() and stream() safely handle being called from within an async context (e.g., Jupyter notebooks, async frameworks).
Multi-turn with messages=¶
run, arun, and astream accept an optional keyword-only messages= — a
list of prior conversation turns inserted after the system prompt + memory
context and before the current input. Default None reproduces the
single-input behavior exactly.
from fastaiagent import Agent
from fastaiagent.llm.message import AssistantMessage, UserMessage
history = [
UserMessage("My name is Alice."),
AssistantMessage("Nice to meet you, Alice!"),
]
result = agent.run("What's my name?", messages=history)
# The model sees the prior turns and can answer "Alice".
This is the building block for Agent Simulation, which drives multi-turn conversations against an agent and judges the transcript. For durable, server-side conversation history across requests, see Memory.
Dynamic Instructions¶
System prompts can be a callable that receives the RunContext, enabling per-request personalization:
agent = Agent(
name="personalized",
system_prompt=lambda ctx: (
f"You help {ctx.state.user_name} with their {ctx.state.plan} plan."
if ctx else
"You are a general-purpose assistant."
),
llm=LLMClient(provider="openai", model="gpt-4.1"),
)
ctx = RunContext(state=UserState(user_name="Alice", plan="pro"))
result = agent.run("Help me", context=ctx)
See Dynamic Instructions for full details.
Serialization¶
Agents can be serialized to JSON and restored:
to_dict() emits {name, agent_type, system_prompt, llm_endpoint, tools, guardrails,
config}, plus two governed fields only when configured (so an agent with neither is
byte-identical to earlier versions):
prompt_slug— set viaAgent(prompt_slug=...); references a control-plane registry prompt. When set,system_promptis emitted as""(the slug is the source of truth).memory_enabled—truewhenmemory=is configured.
This is the canonical payload for pushing an agent to a connected control plane — POST it
to /public/v1/sdk/agents. See Pushing agent definitions.
AgentResult¶
Every agent execution returns an AgentResult. It has twelve fields, and the
table below is all of them:
| Field | Type | Description |
|---|---|---|
output |
str |
The agent's final text response |
parsed |
Any \| None |
The response parsed into your output_type, when the agent declares one. None otherwise — read this, not output, for a structured run. |
tool_calls |
list[dict] |
All tool calls made during execution |
tokens_used |
int |
Tokens across every LLM call the run made — each turn of the tool loop, a structured re-ask, a guardrail re-ask. Same run-scoped accumulator cost is built from, so the two always agree. |
cost |
float |
Estimated USD spend, summed over every LLM call the run made — the tool loop's turns, a structured re-ask, a guardrail re-ask. Priced from the model id and the provider's token counts against the built-in list-price table (override it with set_rate_overrides()). 0.0 for a model with no rate — see cost_known. |
cost_known |
bool |
Whether every call in the run could be priced. cost == 0.0 does not always mean the run was free: a private fine-tune and a bedrock/azure deployment id have no rate, and reporting $0.00 for them would be a guess, so this reads False. A self-hosted provider (ollama, lmstudio, vllm) is the exception — it runs on hardware you already pay for, so it is a known zero and this reads True. Check this flag before trusting the number. |
latency_ms |
int |
Total execution time in milliseconds |
trace_id |
str \| None |
Trace ID for debugging. Populated on every path since 1.67.0 — run, arun, stream, and the same fields on Swarm, Supervisor and Chain. |
execution_id |
str |
The durable execution's id. Always populated when a checkpointer= is configured; "" otherwise. It is what you pass to agent.aresume(...). |
status |
str |
"completed" for a normal run, "paused" when a tool called interrupt(). A paused run returns — it does not raise — so a durability user must branch on this. |
pending_interrupt |
dict \| None |
Set when status == "paused": {reason, context, node_id, agent_path} — the same payload the /approvals UI reads from the pending_interrupts table. |
guardrails |
list[GuardrailFiring] |
Every guardrail that executed, in order, across all four positions. The only way to observe a non-halting outcome: warn, mask and override all let the run finish, so before 1.64.0 a run with the Local UI off and no plane attached reported a clean string whether or not a rule had fired. A blocking failure still raises GuardrailBlockedError rather than returning. |
result = agent.run("my ssn is 123-45-6789")
if result.status == "paused":
print("waiting on", result.pending_interrupt["reason"], result.execution_id)
for g in result.guardrails:
if g.fired():
print(g.name, g.position, g.action_taken)
tokens_used counts the whole run as of 1.68.0 — the number goes up
It used to be read off the last LLM response the tool loop returned, so a
three-call run reported one call's tokens and a re-ask reported none of the
turns that preceded it. It now comes from the run-scoped accumulator
LLMClient writes to at the single point every completion passes through —
the same hook cost uses, which is why the two now agree.
This is a behaviour change with no deprecation window: for any multi-turn run
the reported number increases, and a budget, alert or assertion calibrated
against the old value will see a step change at 1.68.0. Single-call runs are
unaffected. A custom client that implements acomplete without extending
LLMClient never reaches the accumulator and keeps the legacy last-response
sum, because that is a better answer for it than zero.
Where cost comes from, and where it lands
It is computed at the provider call inside LLMClient and also written to
the llm.* span as fastaiagent.cost.total_usd — the same attribute the
LangChain, CrewAI and Pydantic-AI integrations have always emitted, so the
control plane, the Local UI and trace export read one key for every
framework. Before 1.67.0 nothing populated AgentResult.cost and the SDK's
own runs were the only framework the plane received no cost for; the UI hid
it behind a read-time estimate from token counts.
Error Handling¶
from fastaiagent._internal.errors import (
AgentError, # Base agent error
AgentTimeoutError, # Execution timeout
MaxIterationsError, # Tool loop exceeded max_iterations
GuardrailBlockedError, # Guardrail rejected input/output
LLMProviderError, # LLM API error
)
try:
result = agent.run("Do something complex")
except MaxIterationsError:
print("Agent couldn't complete in time")
except GuardrailBlockedError as e:
print(f"Blocked by {e.guardrail_name}: {e}")
except LLMProviderError as e:
print(f"LLM error: {e}")
Middleware¶
Middleware intercepts and transforms messages, responses, and tool calls without subclassing Agent. Use it for message trimming, PII redaction, tool-call budgets, caching, and other cross-cutting concerns.
from fastaiagent import Agent, TrimLongMessages, ToolBudget, RedactPII
agent = Agent(
name="controlled",
llm=LLMClient(provider="openai", model="gpt-4.1"),
middleware=[
TrimLongMessages(keep_last=20),
RedactPII(),
ToolBudget(max_calls=10),
],
)
See Middleware for the full reference, ordering semantics, and how to write your own.
Next Steps¶
- Middleware — Composable pre/post model hooks and tool wrappers
- Dynamic Instructions — Personalize system prompts per-request
- Agent Memory — Give agents conversation memory across turns
- Multi-Agent Teams — Supervisor / Worker (centralized delegation)
- Swarm — Peer-to-peer handoff topology (no coordinator)
- Tools — Deep dive into using tools with agents
- Guardrails — Full guardrail reference
- Chains — Compose agents into multi-step workflows