Skip to content

Agents

An Agent is the core building block of the FastAIAgent SDK. It wraps an LLM with tools, guardrails, and memory to create an autonomous assistant that can reason, take actions, and validate its own output.

Creating an Agent

from fastaiagent import Agent, LLMClient

agent = Agent(
    name="support-bot",
    system_prompt="You are a helpful customer support agent.",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
)

result = agent.run("How do I reset my password?")
print(result.output)

Supported LLM providers:

Provider Example
OpenAI LLMClient(provider="openai", model="gpt-4.1")
Anthropic LLMClient(provider="anthropic", model="claude-sonnet-4-6")
Ollama LLMClient(provider="ollama", model="llama3")
Azure LLMClient(provider="azure", model="gpt-4", base_url="https://myendpoint.openai.azure.com/openai/deployments/gpt-4/")
AWS Bedrock LLMClient(provider="bedrock", model="anthropic.claude-3-sonnet-20240229-v1:0")
Custom LLMClient(provider="custom", model="my-model", base_url="https://my-api.com/v1")

Agent with Tools

Tools let your agent take actions — call APIs, search databases, run calculations.

from fastaiagent import Agent, FunctionTool, LLMClient

def get_weather(city: str) -> str:
    """Get current weather for a city."""
    return f"Sunny, 22°C in {city}"

def search_orders(order_id: str) -> str:
    """Look up an order by ID."""
    return f"Order {order_id}: Shipped, arriving tomorrow."

agent = Agent(
    name="assistant",
    system_prompt="You help customers with weather and order questions. Use tools.",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
    tools=[
        FunctionTool(name="get_weather", fn=get_weather),
        FunctionTool(name="search_orders", fn=search_orders),
    ],
)

result = agent.run("What's the weather in Paris and where is order ORD-123?")
print(result.output)       # LLM's final text response
print(result.tool_calls)   # List of tool calls made
print(result.tokens_used)  # Tokens across EVERY turn of the run, not just the last
print(result.latency_ms)   # Total execution time

How tool calling works: 1. Agent sends messages + tool schemas to the LLM 2. LLM decides to call one or more tools (or respond directly) 3. SDK executes the tools and sends results back to the LLM 4. LLM generates a final response using the tool results 5. This loop repeats up to max_iterations times

The @tool Decorator

For quick tool creation:

from fastaiagent.tool import tool

@tool(name="calculate")
def calculate(expression: str) -> str:
    """Evaluate a math expression."""
    return str(eval(expression))

# Use directly — it's a FunctionTool
result = calculate.execute({"expression": "2 + 2"})

Tool Types

Type Use Case Example
FunctionTool Wrap any Python function FunctionTool(name="calc", fn=my_func)
RESTTool Call an HTTP API RESTTool(name="weather", url="https://api.weather.com/v1", method="GET")
MCPTool Connect to MCP server MCPTool(name="search", server_url="http://localhost:3000")

Agent with Guardrails

Guardrails validate input/output at every step, blocking unsafe content automatically.

from fastaiagent import Agent, LLMClient
from fastaiagent.guardrail import no_pii, toxicity_check, json_valid

agent = Agent(
    name="safe-bot",
    system_prompt="You are a helpful assistant.",
    llm=LLMClient(provider="anthropic", model="claude-sonnet-4-6"),
    guardrails=[
        no_pii(),           # Blocks SSN, email, phone, credit cards in output
        toxicity_check(),   # Blocks toxic language
    ],
)

result = agent.run("What are the benefits of eating healthy?")
print(result.output)  # Clean output passes both guardrails

Important: no_pii() is an output guardrail — it checks the LLM's response, not the user's input. If the LLM happens to include a real SSN, email, or phone number in its response, the guardrail blocks it and raises GuardrailBlockedError. Most LLMs will refuse to output real PII on their own, so this guardrail acts as a safety net for edge cases, tool results that leak PII, or less-guarded models.

To guard against input containing PII, set the position explicitly:

from fastaiagent import Agent, LLMClient
from fastaiagent._internal.errors import GuardrailBlockedError
from fastaiagent.guardrail import no_pii, GuardrailPosition

agent = Agent(
    name="input-safe-bot",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
    guardrails=[
        no_pii(position=GuardrailPosition.input),   # Blocks PII in user input
        no_pii(position=GuardrailPosition.output),   # Blocks PII in LLM output
    ],
)

# User input containing an SSN is blocked before reaching the LLM
try:
    result = agent.run("My SSN is 123-45-6789, can you store it?")
except GuardrailBlockedError as e:
    print(f"Blocked: {e}")  # "PII detected: SSN"

Built-in guardrail factories:

All thirteen are exported from fastaiagent.guardrail. The default position is in the signature — no_pii() is an output guardrail, no_prompt_injection() an input one, allowed_domains() a tool_call one — and every factory takes position= to move it.

Factory What it checks Default position
no_pii(entities=…, backend=…) Email, US phone, SSN, credit cards (Luhn-validated) output
no_secrets() Leaked credentials, API keys and tokens output
no_prompt_injection(mode="heuristic"|"llm") Prompt-injection / jailbreak attempts. on_error="allow" by default input
json_valid() Output is valid JSON output
toxicity_check(mode="keyword"|"llm") Toxic language. on_error="allow" by default output
openai_moderation(model="omni-moderation-latest") Content flagged by the OpenAI moderation endpoint. on_error="block" output
grounded(reference, threshold=0.7) Whether the answer is supported by reference. on_error="block" output
no_hallucination(reference, threshold=0.7) Alias of grounded() — same check, the name you reach for output
banned_topics([...]) Content falling under any banned topic (blocklist). on_error="allow" output
allowed_topics([...]) Content outside the given topics (allowlist gate). on_error="block" output
cost_limit(max_usd=0.10) The run's accumulated LLM spend so far. Blocks over budget; raises (so on_error decides) when the model has no rate in the pricing table, because an unpriced run is not a free one. Until 1.67.0 this always passed. output
allowed_domains(["api.example.com"]) URL domains in tool calls tool_call
responsible_ai(...) Not one guardrail — returns a list to spread into guardrails=[...]. See Responsible AI (several)

See the Guardrails reference for each factory's full signature and on_error for what a degraded check costs.

Custom guardrails:

from fastaiagent.guardrail import Guardrail, GuardrailPosition, GuardrailType

# Inline function
guardrail = Guardrail(
    name="max_length",
    position=GuardrailPosition.output,
    blocking=True,
    fn=lambda text: len(text) < 500,
)

# Regex-based
guardrail = Guardrail(
    name="no_urls",
    guardrail_type=GuardrailType.regex,
    position=GuardrailPosition.output,
    config={"pattern": r"https?://", "should_match": False},
)

Guardrail positions: input, output, tool_call, tool_result

Blocking modes: - blocking=True — raises GuardrailBlockedError if validation fails - blocking=False — logs the failure but continues execution

Agent Configuration

from fastaiagent import Agent, AgentConfig

agent = Agent(
    name="configured-agent",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
    config=AgentConfig(
        max_iterations=5,     # Max tool-calling loop iterations (default: 10)
        tool_choice="auto",   # "auto", "required", "none"
        temperature=0.7,      # LLM temperature override
        max_tokens=1000,      # Max response tokens
    ),
)

Streaming

Stream tokens from the agent as they are generated, rather than waiting for the full response:

from fastaiagent.llm.stream import TextDelta, ToolCallStart

async for event in agent.astream("What's the weather in Paris?"):
    if isinstance(event, TextDelta):
        print(event.text, end="", flush=True)
    elif isinstance(event, ToolCallStart):
        print(f"\n[Calling {event.tool_name}...]")

A sync wrapper is also available:

result = agent.stream("Hello")  # returns AgentResult
print(result.output)

Streaming runs input guardrails before streaming begins and output guardrails after streaming completes. Memory is updated at the end.

See Streaming for full details, event types, and chat UI patterns.

Sync vs Async

Every method has both sync and async versions:

# Sync
result = agent.run("Hello")

# Async
result = await agent.arun("Hello")

Streaming also has both forms:

# Async — yields events in real time
async for event in agent.astream("Hello"):
    ...

# Sync — collects into AgentResult
result = agent.stream("Hello")

The sync run() and stream() safely handle being called from within an async context (e.g., Jupyter notebooks, async frameworks).

Multi-turn with messages=

run, arun, and astream accept an optional keyword-only messages= — a list of prior conversation turns inserted after the system prompt + memory context and before the current input. Default None reproduces the single-input behavior exactly.

from fastaiagent import Agent
from fastaiagent.llm.message import AssistantMessage, UserMessage

history = [
    UserMessage("My name is Alice."),
    AssistantMessage("Nice to meet you, Alice!"),
]
result = agent.run("What's my name?", messages=history)
# The model sees the prior turns and can answer "Alice".

This is the building block for Agent Simulation, which drives multi-turn conversations against an agent and judges the transcript. For durable, server-side conversation history across requests, see Memory.

Dynamic Instructions

System prompts can be a callable that receives the RunContext, enabling per-request personalization:

agent = Agent(
    name="personalized",
    system_prompt=lambda ctx: (
        f"You help {ctx.state.user_name} with their {ctx.state.plan} plan."
        if ctx else
        "You are a general-purpose assistant."
    ),
    llm=LLMClient(provider="openai", model="gpt-4.1"),
)

ctx = RunContext(state=UserState(user_name="Alice", plan="pro"))
result = agent.run("Help me", context=ctx)

See Dynamic Instructions for full details.

Serialization

Agents can be serialized to JSON and restored:

# Serialize
data = agent.to_dict()

# Restore
restored = Agent.from_dict(data)

to_dict() emits {name, agent_type, system_prompt, llm_endpoint, tools, guardrails, config}, plus two governed fields only when configured (so an agent with neither is byte-identical to earlier versions):

  • prompt_slug — set via Agent(prompt_slug=...); references a control-plane registry prompt. When set, system_prompt is emitted as "" (the slug is the source of truth).
  • memory_enabled — true when memory= is configured.

This is the canonical payload for pushing an agent to a connected control plane — POST it to /public/v1/sdk/agents. See Pushing agent definitions.

AgentResult

Every agent execution returns an AgentResult. It has twelve fields, and the table below is all of them:

Field Type Description
output str The agent's final text response
parsed Any \| None The response parsed into your output_type, when the agent declares one. None otherwise — read this, not output, for a structured run.
tool_calls list[dict] All tool calls made during execution
tokens_used int Tokens across every LLM call the run made — each turn of the tool loop, a structured re-ask, a guardrail re-ask. Same run-scoped accumulator cost is built from, so the two always agree.
cost float Estimated USD spend, summed over every LLM call the run made — the tool loop's turns, a structured re-ask, a guardrail re-ask. Priced from the model id and the provider's token counts against the built-in list-price table (override it with set_rate_overrides()). 0.0 for a model with no rate — see cost_known.
cost_known bool Whether every call in the run could be priced. cost == 0.0 does not always mean the run was free: a private fine-tune and a bedrock/azure deployment id have no rate, and reporting $0.00 for them would be a guess, so this reads False. A self-hosted provider (ollama, lmstudio, vllm) is the exception — it runs on hardware you already pay for, so it is a known zero and this reads True. Check this flag before trusting the number.
latency_ms int Total execution time in milliseconds
trace_id str \| None Trace ID for debugging. Populated on every path since 1.67.0 — run, arun, stream, and the same fields on Swarm, Supervisor and Chain.
execution_id str The durable execution's id. Always populated when a checkpointer= is configured; "" otherwise. It is what you pass to agent.aresume(...).
status str "completed" for a normal run, "paused" when a tool called interrupt(). A paused run returns — it does not raise — so a durability user must branch on this.
pending_interrupt dict \| None Set when status == "paused": {reason, context, node_id, agent_path} — the same payload the /approvals UI reads from the pending_interrupts table.
guardrails list[GuardrailFiring] Every guardrail that executed, in order, across all four positions. The only way to observe a non-halting outcome: warn, mask and override all let the run finish, so before 1.64.0 a run with the Local UI off and no plane attached reported a clean string whether or not a rule had fired. A blocking failure still raises GuardrailBlockedError rather than returning.
result = agent.run("my ssn is 123-45-6789")

if result.status == "paused":
    print("waiting on", result.pending_interrupt["reason"], result.execution_id)

for g in result.guardrails:
    if g.fired():
        print(g.name, g.position, g.action_taken)

tokens_used counts the whole run as of 1.68.0 — the number goes up

It used to be read off the last LLM response the tool loop returned, so a three-call run reported one call's tokens and a re-ask reported none of the turns that preceded it. It now comes from the run-scoped accumulator LLMClient writes to at the single point every completion passes through — the same hook cost uses, which is why the two now agree.

This is a behaviour change with no deprecation window: for any multi-turn run the reported number increases, and a budget, alert or assertion calibrated against the old value will see a step change at 1.68.0. Single-call runs are unaffected. A custom client that implements acomplete without extending LLMClient never reaches the accumulator and keeps the legacy last-response sum, because that is a better answer for it than zero.

Where cost comes from, and where it lands

It is computed at the provider call inside LLMClient and also written to the llm.* span as fastaiagent.cost.total_usd — the same attribute the LangChain, CrewAI and Pydantic-AI integrations have always emitted, so the control plane, the Local UI and trace export read one key for every framework. Before 1.67.0 nothing populated AgentResult.cost and the SDK's own runs were the only framework the plane received no cost for; the UI hid it behind a read-time estimate from token counts.

Error Handling

from fastaiagent._internal.errors import (
    AgentError,              # Base agent error
    AgentTimeoutError,       # Execution timeout
    MaxIterationsError,      # Tool loop exceeded max_iterations
    GuardrailBlockedError,   # Guardrail rejected input/output
    LLMProviderError,        # LLM API error
)

try:
    result = agent.run("Do something complex")
except MaxIterationsError:
    print("Agent couldn't complete in time")
except GuardrailBlockedError as e:
    print(f"Blocked by {e.guardrail_name}: {e}")
except LLMProviderError as e:
    print(f"LLM error: {e}")

Middleware

Middleware intercepts and transforms messages, responses, and tool calls without subclassing Agent. Use it for message trimming, PII redaction, tool-call budgets, caching, and other cross-cutting concerns.

from fastaiagent import Agent, TrimLongMessages, ToolBudget, RedactPII

agent = Agent(
    name="controlled",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
    middleware=[
        TrimLongMessages(keep_last=20),
        RedactPII(),
        ToolBudget(max_calls=10),
    ],
)

See Middleware for the full reference, ordering semantics, and how to write your own.

Next Steps

  • Middleware — Composable pre/post model hooks and tool wrappers
  • Dynamic Instructions — Personalize system prompts per-request
  • Agent Memory — Give agents conversation memory across turns
  • Multi-Agent Teams — Supervisor / Worker (centralized delegation)
  • Swarm — Peer-to-peer handoff topology (no coordinator)
  • Tools — Deep dive into using tools with agents
  • Guardrails — Full guardrail reference
  • Chains — Compose agents into multi-step workflows