Agent Memory¶
Memory lets agents remember — within a conversation and across sessions. The recommended API is a single object, Memory, with progressive-disclosure keywords. The composable blocks it's built on remain available for advanced/custom behaviours (see Advanced).
New to memory here? Start with How memory works — the kinds of memory, tiers vs scopes, what is written when, and the two "learn"s.
Memory — the recommended API¶
from dataclasses import dataclass
from fastaiagent import Agent, LLMClient, Memory, RunContext
llm = LLMClient(provider="openai", model="gpt-4.1")
# Just remember the conversation:
agent = Agent(name="assistant", llm=llm, memory=Memory())
# Personalize per user, multi-user safe, and learn durable facts:
agent = Agent(name="support", llm=llm, memory=Memory(
user_id=lambda ctx: ctx.state.user_id, # resolved per run — one agent, many users
learn=llm, # extract + persist durable user facts
))
@dataclass
class Session:
user_id: str
# Each run names its user through the run context the resolver receives.
agent.run("What's my plan?", context=RunContext(state=Session(user_id="alice")))
The whole surface is keywords on one object:
| Keyword | What it does |
|---|---|
location |
where durable facts live — "sqlite" (default), postgres://… / redis://…, or a store instance |
user_id |
personalization key — a string or (ctx)->str resolver (per-run, multi-user) |
agent_id |
global tier: facts true for everyone using the agent |
project_id |
tenant partition applied across tiers |
window |
recent messages kept (session/working memory) |
learn |
an LLM → extract + persist durable facts from each user message (never from the model's replies) |
max_learned_facts |
keep the newest N learned facts per user, deleting older ones (default 200; None = no cap). Facts you persist yourself are never touched |
summarize |
an LLM → compress older turns into a running summary |
recall |
"auto" (an in-process FAISS index per user) or a VectorStore shared by every user, each user's recall namespaced → semantic recall of past exchanges |
dedupe |
drop recalled content an earlier tier already injected |
semantic |
"auto" or a VectorStore → retrieve(query, ...) by meaning |
embedder |
the embedder for semantic and recall; "auto" indexes are sized to it |
max_users |
with a user_id resolver, keep at most N users' windows in the process, dropping the least recently used (default 10_000; None = no cap). A dropped user's durable facts stay |
plane_agent_id |
the agent's id on a connected Enterprise plane → also inject the plane's curated facts, for every user (see PlaneFactBlock) |
Tiers — who a fact is true for¶
global→ true for everyone using the agent (store scopeagent).user→ per-user personalization (store scopeuser; needs an id).session→ the ephemeral conversation window (window) — not a durable store.
Direct store use¶
mem = Memory(location="sqlite", agent_id="support")
mem.persist("Return policy is 30 days", tier="global") # create; returns fact id
mem.persist("Prefers email", tier="user", id="alice")
mem.retrieve(tier="user", id="alice") # read → list[Fact]
mem.update("Prefers Slack", old="Prefers email", tier="user", id="alice") # supersede old, keep history
mem.forget(tier="user", id="alice") # delete; returns count
A global fact is filed under the Memory's agent_id, and only a Memory(agent_id=...) with the same id injects it. persist/update with tier="global" and no agent_id warn: that fact would never be injected.
forget refuses to mass-delete by accident. forget(tier="user") needs an id, and forget(tier="global") needs agent_id= on the Memory (or an id) — without one, an empty id would match every agent's global facts. Pass id="*" to delete every subject on purpose.
Facts are versioned by supersede, never overwritten — update marks the old row superseded (kept in the audit history, visible under the Memory page's "Show superseded" toggle) and activates the new one. forget hard-deletes (including superseded history for that subject).
Multi-user safety (important)¶
Memory(user_id=<resolver>) gives each user their own working memory — durable facts and the live session window — keyed on the resolved id. One agent definition safely serves many users, with run, arun and astream alike. The resolver receives the run's context=, so pass the user there on every run (see the example above).
A caller the resolver can't resolve gets no conversation memory. That covers a run with no context=, a resolver that returns None or "", and a resolver that raises. It sees global facts only, and nothing it says is kept, so unresolved callers can never share a window. A raising resolver logs one warning. The usual cause is a dict state read as an attribute: use lambda ctx: ctx.state["user_id"] for RunContext(state={"user_id": ...}).
Outside a run there is no current user, so save/load on a per-user Memory raise. Use for_user instead:
mem.for_user("alice").save("memory/alice") # persist Alice's window
mem.for_user("alice").load("memory/alice") # restore it before her next run
Windows live in the process:
- Per-user working windows are held in-process. Only durable facts move to an external
location; for horizontally-scaled deployments, persist windows withfor_user(...).save(...)or keep sessions sticky. recall="auto"builds a per-user in-process vector store. AVectorStoreyou pass is shared by every user, and each user's recall is namespaced (user:<id>, prefixed withproject_idwhen set), so one user's messages never come back for another.
How many users' windows are kept (max_users)¶
A long-running server can see far more users than it needs to keep in memory at once. Memory keeps the windows of at most max_users users (default 10_000) and, past that, drops the one used least recently:
mem = Memory(
location="postgres://user:pw@host:5432/db",
user_id=lambda ctx: ctx.state.user_id,
max_users=50_000, # or None for no cap (memory grows with every new user)
)
What a dropped user loses and keeps:
| Dropped | Kept |
|---|---|
the conversation window, the running summary, the recall="auto" index |
every durable fact — they live in the store and are read again on the user's next turn |
So a returning user starts a fresh conversation but is still recognised — runnable in examples/100_memory_in_production.py. The first drop logs one warning. Nothing is saved for you when a window is dropped: to carry a user's conversation across a drop (or a restart), save it with mem.for_user(id).save(path) and load it before their next run.
In the rare case that more than max_users other users become active while one user's run is in flight, that run's reply can land in a fresh window. Size max_users well above your peak number of concurrent users.
Storage backends (location)¶
Durable facts live wherever location points — the same Memory API, a different backend:
Memory(location="sqlite") # default local.db (single-node)
Memory(location="postgres://user:pw@host:5432/db") # needs fastaiagent[postgres]
Memory(location="redis://host:6379/0") # needs fastaiagent[redis]
Memory(location=my_store) # any object implementing FactStore
All backends implement the same FactStore contract (idempotent add — also under concurrency: writers that add the same fact at once all get the one row's id — safe scoping, supersede, delete, newest-first reads) and are verified against one shared conformance suite, so behaviour doesn't drift. Use Postgres/Redis for multi-node or many-user deployments; SQLite for dev/single-node. Runnable demo: examples/memory_backends/.
Every turn reads the newest facts for the user (and the agent), so a read costs what it returns, not what is stored: Postgres reads through partial indexes, and Redis through newest-first sorted indexes.
Connections are reused. The Postgres store keeps a pool of 1–10 connections per process (a store created before a fork, as under gunicorn --preload, opens its own pool in each worker). The pool is closed when the store is garbage-collected and at interpreter exit. To size it or release it yourself, build the store:
from fastaiagent.learn import PostgresFactStore
store = PostgresFactStore("postgres://user:pw@host:5432/db", min_pool_size=1, max_pool_size=10)
mem = Memory(location=store)
...
store.close() # releases the connections; the store reopens on its next use
Upgrading a Redis or Postgres store to 1.80
Redis. The first time a store opens a namespace written by an older SDK, it indexes the existing facts once (a SCAN over the namespace). If an older SDK keeps writing to the same namespace afterwards, its new facts are missing from reads until you call RedisFactStore(url).reindex() — upgrade every writer together.
Postgres. The first open creates two indexes on learned_memory. On a very large table, create them yourself beforehand so the build doesn't block writes; the store then finds them and skips the step:
CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_learned_memory_subject_active
ON learned_memory (scope, scope_id, project_id, created_at DESC)
WHERE superseded_by IS NULL;
CREATE INDEX CONCURRENTLY IF NOT EXISTS idx_learned_memory_scope_active
ON learned_memory (scope, project_id, created_at DESC)
WHERE superseded_by IS NULL;
Observability with external backends
Agent runs against any backend emit memory.read / memory.write trace spans, and direct Memory.persist / retrieve / update / forget calls emit memory.persist / memory.retrieve / … spans (browsable in fastaiagent ui). The UI Memory page browses the local SQLite store; facts written to an external backend are observed via those trace spans (a shared UI over external backends is future work).
Semantic recall of facts (semantic)¶
By default retrieve(tier=, id=) returns a scope's facts as a list. Turn on semantic= to retrieve by meaning with retrieve(query, tier=, id=):
mem = Memory(location="sqlite", semantic="auto") # in-process FAISS + default embedder
mem.persist("The user is allergic to peanuts", tier="user", id="alice")
mem.retrieve("what foods should we avoid?", tier="user", id="alice") # → the peanut fact
semantic="auto" builds an in-process vector index sized to the embedder; pass a VectorStore (FAISS, Qdrant, Chroma, your own) for a shared/production index and embedder= to override. Facts written by learn= are indexed automatically (they share the store). Semantic results honor the same scope isolation and skip superseded facts; the memory.retrieve span records how many facts came back.
The store is the source of truth, not the index. Before each semantic search, the subject's active facts are read from the store and any the index doesn't have yet are embedded, in one batch. So a restarted process (whose in-process index starts empty) and facts written by another process are both found; the first query after a restart pays one embedding call for that subject's facts. Each fact's vector id is a stable UUID with the fact id in its metadata, which Qdrant requires.
Safe-by-default scoping¶
At user/project scope an empty id returns nothing — one user's facts can never leak into another's context. Use scope_id="*" (on the low-level store) to deliberately read across all subjects. The agent/global tier stays permissive (shared truth), but an empty agent id is refused where it is almost always a mistake: Memory(agent_id="") raises, and PersistentFactBlock(scope="agent", scope_id="") warns — use scope_id="*" to read every agent's facts on purpose. (This corrects prior behaviour where an empty user id matched everyone — see the CHANGELOG.)
The Memory page (fastaiagent ui → Knowledge → Memory) shows the tiers side by side — user:alice / user:bob (learned, source trace) and the global fact under agent:assistant (source manual, from Memory(agent_id="assistant")):

Reproduce with examples/memory_simple/ (companion.py runs one agent for two users; snapshot.py captures the UI).
Advanced: composable blocks¶
The block API below powers
Memoryunder the hood and remains fully supported for custom behaviours (write your ownMemoryBlock, control ordering, etc.). Most apps only needMemoryabove.
AgentMemory |
ComposableMemory |
|
|---|---|---|
| Since | 0.1.x | 0.4.0 |
| Stores | Sliding window of raw messages | Sliding window plus any number of long-term blocks |
| Best for | Chatbots, short sessions | Long-running assistants, personal memory, fact tracking |
| Drop-in replacement | — | Yes — Agent(memory=...) accepts either |
Sliding-window memory: AgentMemory¶
from fastaiagent import Agent, LLMClient
from fastaiagent.agent import AgentMemory
memory = AgentMemory(max_messages=20)
agent = Agent(
name="assistant",
system_prompt="Remember what users tell you. Be brief.",
llm=LLMClient(provider="openai", model="gpt-4.1"),
memory=memory,
)
agent.run("My name is Alice.")
result = agent.run("What's my name?")
print(result.output) # "Your name is Alice."
How it works¶
- On each
run()call, the agent prepends stored messages to the conversation. - After the agent responds, the new user message and assistant response are added to memory. A resumed or forked run records the question it was resuming. The stored reply is the final answer after middleware (with
RedactPII, redacted) — only the last turn's reply, not text the model said before calling a tool — andrunandastreamstore the same. The user's message is stored as they said it. - If
max_messagesis reached, the oldest messages are dropped (FIFO).
Persistence¶
memory.save("memory.json")
new_memory = AgentMemory()
new_memory.load("memory.json")
agent = Agent(name="assistant", llm=..., memory=new_memory)
result = agent.run("What's my name?") # "Your name is Alice."
Configuration¶
| Parameter | Type | Default | Description |
|---|---|---|---|
max_messages |
int \| None |
None |
Messages to retain (a count, not tokens). None keeps everything; Memory(window=) defaults to 20 |
Long-term memory: ComposableMemory + blocks¶
ComposableMemory wraps a primary AgentMemory sliding window with a list of memory blocks that contribute SystemMessage fragments to every turn:
┌───────────────────────────────────────────────┐
│ ComposableMemory.get_context(query) │
│ │
│ ← SystemMessage(s) from block[0].render │
│ ← SystemMessage(s) from block[1].render │
│ ← SystemMessage(s) from block[n].render │
│ ← primary window (last N messages) │
└───────────────────────────────────────────────┘
Quick start¶
from fastaiagent import Agent, LLMClient
from fastaiagent.agent import (
ComposableMemory, AgentMemory,
StaticBlock, SummaryBlock, VectorBlock, FactExtractionBlock,
)
from fastaiagent.kb.backends.faiss import FaissVectorStore
llm = LLMClient(provider="openai", model="gpt-4o-mini")
memory = ComposableMemory(
blocks=[
StaticBlock("The user's name is Upendra and they prefer terse answers."),
SummaryBlock(llm=llm, keep_last=10, summarize_every=5),
VectorBlock(store=FaissVectorStore(dimension=384, index_type="flat")),
FactExtractionBlock(llm=llm, max_facts=100),
],
primary=AgentMemory(max_messages=20),
)
agent = Agent(name="assistant", llm=llm, memory=memory)
Each block is optional and independently useful. Use only what you need.
Built-in blocks¶
StaticBlock¶
A fixed system-level fact, injected on every turn. Zero state, zero LLM calls.
SummaryBlock¶
Maintains a rolling LLM-generated summary of older turns. Refreshes every summarize_every messages, summarizing everything older than keep_last.
SummaryBlock(
llm=llm,
keep_last=10, # never summarize the N most recent messages
summarize_every=5, # refresh cadence
max_chars=800, # asked of the LLM, and longer summaries are cut here
)
When to use: long conversations that otherwise blow the context window. Cheaper than re-embedding everything, but introduces one extra LLM call every summarize_every messages (a turn is two: the user's and the answer).
VectorBlock¶
Semantic recall over past messages. Each incoming message (above min_content_chars) is embedded and stored in a VectorStore. On each turn, the query is embedded and the top-k most similar past messages are surfaced.
from fastaiagent.kb.backends.faiss import FaissVectorStore
VectorBlock(
store=FaissVectorStore(dimension=384, index_type="flat"),
top_k=5,
namespace="default", # tag so multiple blocks can share a store
min_content_chars=10, # skip trivial messages ("ok", "yes")
)
Any backend implementing the VectorStore protocol works — FaissVectorStore, QdrantVectorStore, ChromaVectorStore, your own. See KB Backends.
Blocks sharing one store recall only their own namespace. The search over-fetches and widens until it has top_k of the block's own chunks, so a busy namespace can't crowd out a quiet one. Chunks with no namespace tag (for example, a store shared with a knowledge base) are read by the "default" namespace only.
When to use: conversations that span days or sessions, where long-ago facts should be retrievable by meaning, not just recency.
VectorBlock also accepts dedupe_against_upstream=True — see Shared memory context — to skip recalling anything an earlier block already put in the prompt.
Memory scoring (recency + importance) — v1.9.0¶
By default VectorBlock ranks retrieval results by cosine similarity
alone. That works fine for short sessions but breaks in two recurring
ways for long-running agents:
- Old correct answer drowns out new correct answer. The user said in turn 5 "my email is alice@old.com". In turn 50 they corrected themselves: "actually, alice@new.com". Both messages are about email — both score high on similarity to "what's my email?". The older one wins half the time because there's nothing differentiating them.
- Trivial chatter outranks load-bearing facts. The user said "I'm a vegetarian" once. They've then said "ok", "thanks", "yes" twenty times. A query like "what should I order for dinner?" has higher similarity to the "ok" / "thanks" cluster than to that one important fact, and the agent forgets the constraint.
Three optional knobs fix both. Defaults are zero, so existing
VectorBlock(store=...) calls keep their current behaviour byte-for-byte:
VectorBlock(
store=...,
top_k=5,
recency_weight=0.3, # 0.0–1.0
importance_weight=0.2, # 0.0–1.0
recency_half_life_seconds=3600.0, # 1 hour: recency halves every hour
)
Retrieval becomes a weighted sum of three signals:
final_score = (1 - recency_weight - importance_weight) * cosine_similarity
+ recency_weight * 0.5 ** (age_seconds / half_life)
+ importance_weight * importance
cosine_similarity— what's there today. Range ~0–1.recency— exponential decay from the chunk'screated_at. Withhalf_life=3600s, a message 1 hour old contributes 0.5; 2 hours old, 0.25; a week old, ~0.importance— read from the chunk'smetadata['importance'](default1.0if not set). ForPersistentFactBlock, this is sourced from theconfidencecolumn onlearned_memory, so facts the LLM extracted with high confidence outrank uncertain ones.
Worked example¶
Three stored messages, all matching "what's my email?":
| Message | similarity | age | importance |
|---|---|---|---|
| A: "my email is alice@old.com" | 0.85 | 7 days | 0.5 (likely superseded) |
| B: "actually it's alice@new.com" | 0.80 | 1 hour | 1.0 |
| C: "thanks!" | 0.30 | 5 min | 1.0 |
Today (similarity-only, both new weights = 0): A wins (0.85 > 0.80 > 0.30). Wrong answer.
With recency_weight=0.3, importance_weight=0.2 and a 1-hour
half-life:
A: 0.5*0.85 + 0.3*0.5**(604800/3600) + 0.2*0.5 ≈ 0.425 + ~0.000 + 0.10 = 0.525
B: 0.5*0.80 + 0.3*0.5**(3600/3600) + 0.2*1.0 ≈ 0.400 + 0.150 + 0.20 = 0.750
C: 0.5*0.30 + 0.3*0.5**(300/3600) + 0.2*1.0 ≈ 0.150 + 0.283 + 0.20 = 0.633
B wins — the right answer surfaces. C ranks high on recency but its low similarity prevents it from outranking B.
Tuning¶
- Customer-support bot, current state matters most:
recency_weight=0.4,recency_half_life_seconds=1800(30 min). The most recent user message almost always reflects current intent. - Research assistant, old facts still relevant:
recency_weight=0,importance_weight=0.3. Don't decay; let importance differentiate. - Long-running personal assistant:
recency_weight=0.2,importance_weight=0.3,recency_half_life_seconds=86400(1 day). Slow decay, importance-aware.
Where importance comes from¶
VectorBlockreadschunk.metadata['importance'], defaulting to1.0. Messages the agent records carry no importance (Messagehas no such field), so they all score1.0. To weight chunks, write them into the store yourself withmetadata={"namespace": ..., "importance": 0.3}, or subclassVectorBlockand override_make_chunkto set it.PersistentFactBlockreads theconfidencecolumn onlearned_memoryrows:0.6for facts learned during a run (Memory(learn=),FactExtractionBlock(persist=True)),1.0for facts fromfastaiagent learnand for facts you write yourself unless you set it. Withimportance_weight > 0, run-learned facts rank below the rest.
Backward compatibility: with both weights at zero (the default), behaviour is byte-identical to v1.8.x — the scorer short-circuits and returns the input order unchanged.
FactExtractionBlock¶
Uses a cheap LLM to extract durable facts from each user/assistant message and stores them as a dedup'd list. Renders as a bullet-point Known facts: … SystemMessage.
FactExtractionBlock(
llm=llm, # use a fast model (gpt-4o-mini, claude-haiku)
max_facts=200, # cap; oldest drop on overflow
extract_every=1, # run extraction every N inspected messages
roles=("user", "assistant"), # which messages to read; ("user",) skips the model's own claims
inject=True, # False = extract (and persist) without rendering
)
roles=("user",) keeps the model's replies from being recorded as facts about the user, and halves the extraction calls. Set inject=False when a PersistentFactBlock already reads the same store, so each fact reaches the prompt once — Memory(learn=) does both.
When to use: user-focused assistants where you want stable facts ("user is allergic to peanuts", "user's kids are named Maya and Omar") to persist independently from the conversation log.
Persisting facts across runs (persist=True)¶
By default FactExtractionBlock holds facts only for the current conversation. Set persist=True to also write each newly extracted fact to the durable learned_memory table during the run — so it survives restarts and can be read back later by PersistentFactBlock. This closes the loop without needing the offline fastaiagent learn job.
FactExtractionBlock(
llm=llm,
persist=True, # write new facts to learned_memory during the run
scope="user", # 'user' | 'project' | 'agent'
scope_id="upendra", # REQUIRED when persist=True
confidence=0.6, # stamped on auto-facts; below curated 1.0 so they sort lower
max_persisted=200, # keep the newest N learned facts per subject; None = no cap
)
max_persisted deletes the oldest learned facts beyond the cap after each write. Only facts learned from a run (those with a source_trace_id) count and are deleted; facts written directly are never touched. Deleted facts are gone, not superseded.
Each persisted fact is stamped with the current trace id as source_trace_id, so the Memory page shows a clickable link back to the run that produced it. Writes are idempotent (the store's uniqueness constraint dedupes) and failure-isolated (a store error logs and the run continues). Because it now writes an external store mid-run, isolated_copy() raises MemoryIsolationError when persist=True — the same guard as VectorBlock — so fastaiagent.optimize candidates don't bleed writes.
This is the runtime-write counterpart to the read-only PersistentFactBlock: extract-and-persist on the way in, read-back on future runs. You can still write facts directly with MemoryStore.add(Fact(...)) (source stays NULL / "manual").
PersistentFactBlock¶
Read-only block that loads facts from the learned_memory table — populated offline by fastaiagent learn, the Trace Learning Loop. Carries durable facts across runs, where FactExtractionBlock only carries them within a single conversation.
from fastaiagent.agent.memory_blocks import PersistentFactBlock
PersistentFactBlock(
scope="agent", # 'user' | 'project' | 'agent'
scope_id="my-agent", # identifier within scope; "*" = every agent
project_id="", # optional project filter
max_facts=50, # newest-first cap
refresh_every=1, # re-query store every N renders (1 = always)
# v1.9.0: optional scoring on top of `list_active`'s newest-first order.
# Defaults are 0.0 — historical behaviour preserved.
recency_weight=0.0, # 0.0–1.0
importance_weight=0.0, # 0.0–1.0; importance ← learned_memory.confidence
recency_half_life_seconds=86400.0, # 1 day; facts decay slower than messages
)
PersistentFactBlock honours the same recency_weight /
importance_weight model documented under VectorBlock →
Memory scoring. Importance is sourced from the
existing confidence column on learned_memory.
When to use: long-running agents that should accumulate operational knowledge across sessions ("the writer should always cite token-cost claims", "this user prefers reports under 800 words"). Pair with fastaiagent learn to populate the underlying table from past traces.
Pairs with: FactExtractionBlock for the in-conversation extraction; PersistentFactBlock for the cross-conversation re-injection. Both can live in the same ComposableMemory.
PlaneFactBlock (connected central memory)¶
Read-only block that reads curated, human-approved facts from a connected Enterprise plane via GET /public/v1/memory/facts, and injects them at the start of each turn. Where PersistentFactBlock reads facts from the local learned_memory table, PlaneFactBlock reads the governed/curated facts the plane serves — the read side of central governed memory. The plane extracts durable facts from already-ingested traces and a human curates them; the SDK only reads (there is no SDK fact-push path).
With Memory, one keyword adds it — every user keeps their own window, and callers with no user still get the plane's facts:
import fastaiagent as fa
fa.connect(api_key="fa-...", target="https://your-plane.example.com")
agent = fa.Agent(
name="support",
llm=llm,
memory=fa.Memory(
agent_id="support", # local global facts (optional)
user_id=lambda ctx: ctx.state.user_id, # one window per user
plane_agent_id="my-agent-id", # + the plane's curated facts
),
)
To tune the block (category filter, max_facts, refresh_every, …), compose it yourself:
import fastaiagent as fa
from fastaiagent import Agent, AgentMemory, ComposableMemory
from fastaiagent.agent.memory_blocks import PlaneFactBlock, PersistentFactBlock
fa.connect(api_key="fa-...", target="https://your-plane.example.com")
memory = ComposableMemory(
primary=AgentMemory(),
blocks=[
PlaneFactBlock(
agent_id="my-agent-id", # the agent's id on the plane (required)
category=None, # optional category filter
max_facts=50, # cap per turn (1..200)
query_conditioned=True, # pass the user input for semantic recall
score_threshold=0.0, # min similarity for query-conditioned recall
refresh_every=1, # re-read the plane every N renders (raise to cache)
),
PersistentFactBlock(scope="agent", scope_id="my-agent"), # local facts too (optional)
],
)
agent = Agent(name="support", system_prompt="...", llm=llm, memory=memory)
Read-only and degradable. When the SDK is not connected, or the plane answers that this agent has no facts for this key (e.g. 403: the domain isn't entitled; 404: unknown agent), PlaneFactBlock injects nothing and the agent runs normally — central facts are an enhancement, never a dependency; each such status is logged once per agent, not once per user. If the plane is down — unreachable, timing out, or answering 502 / 503 / 504 from the proxy in front of it (what a production plane returns while its backend restarts), or 429 — the block keeps serving the facts it last fetched and leaves the plane alone for 30 seconds (PlaneFactBlock.retry_after_seconds), or for as long as a Retry-After header asks (up to 10 minutes), before trying again — so an outage costs one wait and one warning, not one per turn. The pause is shared by every block reading the same plane and agent, so with Memory(plane_agent_id=...) one user's failed read spares every other user the wait; when the pause ends, one turn retries. It logs once when the plane is back. The read is a bounded start-of-run network GET (like VectorBlock's search), cached per refresh_every — but a new question always refetches, so one question's facts are never served for another. It never pushes anything. The plane runs no agent code — it serves facts; recall and injection happen locally.
Each fact is injected once. The plane can serve the same fact more than once — two approved copies of it, or one entry per matching vector on a semantic read. The block keeps the first copy (the plane's most important) and drops the rest, comparing text without regard to case or spacing. The trace records how many were dropped, as memory.deduped_count on the memory.read.plane_facts span. The duplicates themselves stay on the plane; remove them on the Agent Memories page. Facts the plane serves are not compared with those from other blocks, so a fact kept both locally and on the plane appears under both headings.
Your users' questions stay home when payloads are off. With query_conditioned=True the user's question is sent to the plane for semantic recall. When FASTAIAGENT_TRACE_PAYLOADS=0, it is left out: the plane then returns the agent's facts by importance rather than by relevance to the question.
When to use: connected (Enterprise) deployments that want a single governed, curated knowledge base shared across a fleet of agents, with central redaction / right-to-be-forgotten. See Connected central memory and the memory loop.
Pairs with: PersistentFactBlock (local facts) — Memory(agent_id=..., plane_agent_id=...) merges local and central knowledge per user; a hand-built ComposableMemory has a single window shared by everyone who uses it.
The curated facts the block reads are managed on the plane's Agent Memories page — created or approved by a human (or extracted from traces), with redaction / right-to-be-forgotten:

A runnable end-to-end example is in examples/87_connected_memory.py — the block on its own, then Memory(plane_agent_id=...) for two users.
Memory on a pushed agent¶
When you push an agent definition to a connected control plane
(Agent.to_dict() → POST /public/v1/sdk/agents, see
Platform), a configured memory=
surfaces as memory_enabled: true in the payload, so the console shows memory Enabled
for that agent. The flag reflects only that memory is configured — the block-level tuning
(what facts, retrieval count, thresholds) stays in your SDK code. An agent with no memory=
omits the field entirely.
Composing blocks¶
Block order matters — they render in declaration order, and the resulting SystemMessages appear in the prompt in that order. Typical ordering:
StaticBlock— hard facts that never changeSummaryBlock— what has happened so farFactExtractionBlock— what we know about the userVectorBlock— relevant past exchanges
Followed by the primary sliding window's recent messages.
Shared memory context¶
Blocks render in declaration order, and each block can optionally read what the earlier blocks already produced this turn. ComposableMemory passes a SharedMemoryContext down the chain — a minimal one-directional pipe (not a full graph): block N sees the output of blocks 1..N-1, never the reverse.
Sharing is opt-in and backward-compatible. The base method delegates to render, so blocks that only implement render(query) — including custom and third-party blocks — are unaffected:
class MemoryBlock:
def render_with_context(self, query, shared):
return self.render(query) # default: ignore upstream
A block that wants upstream output overrides render_with_context and reads from shared:
shared.upstream_text() # concatenated content of all earlier blocks
shared.by_block("static") # messages from a specific earlier block
shared.upstream_messages() # everything upstream, in order
Shipped consumer — VectorBlock(dedupe_against_upstream=True). When enabled, VectorBlock drops any recalled message whose content an earlier block already injected (e.g. a StaticBlock pin or an extracted fact), so you don't spend tokens saying the same thing twice. Matching is a conservative normalized substring test — it only drops clear duplicates. The memory.read.vector span reports deduped_count so you can see how many were skipped. Off by default; VectorBlock behaviour is unchanged unless you set the flag.
Only block output flows through the pipe — the raw primary window is appended afterwards and is not shared.
Persistence¶
ComposableMemory.save(path) writes to a directory:
path/
├── primary.json # sliding window
└── blocks/
├── summary.json # SummaryBlock state
├── facts.json # FactExtractionBlock state
└── static.json # (no-op for StaticBlock; file omitted)
load(path) restores into the same blocks, matched by block.name. You must reconstruct the blocks (with the same llm / store / embedder) before calling load — blocks that hold live resources (LLM clients, vector stores) are not themselves serialized.
memory.save("/var/state/agent-alice")
# ... later, new process ...
memory = ComposableMemory(blocks=[SummaryBlock(llm=...), FactExtractionBlock(llm=...)])
memory.load("/var/state/agent-alice")
Writing your own block¶
Subclass MemoryBlock and implement on_message and render:
from fastaiagent.agent import MemoryBlock
from fastaiagent.llm.message import Message, SystemMessage
class MoodBlock(MemoryBlock):
"""Tracks the user's emoji reactions and pins the latest mood."""
name = "mood"
def __init__(self):
self.latest_mood = ""
def on_message(self, message: Message) -> None:
content = message.content or ""
for emoji in ("🎉", "😡", "😊", "😢"):
if emoji in content:
self.latest_mood = emoji
def render(self, query: str):
if not self.latest_mood:
return []
return [SystemMessage(f"User's latest mood: {self.latest_mood}")]
Then just drop it into ComposableMemory(blocks=[MoodBlock(), ...]).
If your block holds persistent state worth saving, override save(path) and load(path). See fastaiagent/agent/memory_blocks.py for the shipped implementations.
Observability — seeing what the agent remembered¶
Memory is no longer a black box at runtime. Every turn, the read and write are wrapped in trace spans with a child span per block, so you can open a trace and see which block recalled what, and why — the same way you already read KB retrieval.* spans.
memory.read(one per turn) —memory.block_count,memory.message_count, and thememory.query. Onememory.read.<block>child per rendering block carriesmemory.rendered_count, boundedmemory.snippetsof what it injected, and — forVectorBlock—memory.scores(per-item similarity in rank order) plusmemory.deduped_countwhen upstream dedupe is on. This is the difference between "memory recalled something" and "memory recalled the wrong thing, score 0.71".memory.write(one per stored message) — onememory.write.<block>child per block with amemory.action(embedded,summarized,extracted_facts,stored,noop) and amemory.detailcount (e.g.facts_extracted, andpersistedwhenpersist=True).

Click the memory.read.vector child to see the recalled items and their scores:

These spans nest under the agent span automatically and are no-ops when tracing is off — memory behaves exactly as before, with no extra embedding or LLM calls. Snippets, query text and the user id (memory.scope_id) are dropped from exported spans when FASTAIAGENT_TRACE_PAYLOADS=0 (the local trace store keeps them), and honor any installed RedactionPolicy (the "Mask secrets" toggle), since memory content can contain PII. Inside Memory, the global fact block's spans are memory.read.persistent_facts and the user fact block's are memory.read.persistent_facts.user.
The Memory page¶
The Local UI (fastaiagent ui) has a Memory page (sidebar → Knowledge → Memory) that browses the learned_memory table — the durable facts PersistentFactBlock reads back across runs. Filter by scope, trace a fact's source, and toggle masking:

Each row is one durable fact:
| Column | Meaning |
|---|---|
| Fact | The stored statement, e.g. "Has a beagle named Biscuit; allergic to cats." This is the fact text a PersistentFactBlock injects into a matching agent's prompt. |
| Scope | Rendered as scope:scope_id (e.g. user:upendra). scope is one of user / project / agent; scope_id is the identifier within it (a user id, a project key, or an agent name). A block reading PersistentFactBlock(scope="user", scope_id="upendra") will pick up exactly the user:upendra rows. |
| Source | Where the fact came from. A trace link jumps to the run that produced it (facts persisted by FactExtractionBlock(persist=True) carry the run's trace id as source_trace_id). manual = inserted directly via MemoryStore.add. |
| Confidence | The confidence column (0–1); also drives importance_weight ranking. Facts learned during a run (Memory(learn=), FactExtractionBlock(persist=True)) get 0.6; facts from fastaiagent learn and direct writes get 1.0 unless you set it. |
| Created | When the fact was written. |
Scope filter — the dropdown lists every scope:scope_id partition (users, projects, and agents — memory is not agent-only) with counts, plus "All scopes".
History — toggle Show superseded to include facts that a newer version replaced. Superseded rows render muted with a superseded → #id marker pointing at the row that replaced them (facts are versioned by append + supersede(old, new), never overwritten).

Where do these rows come from? Three ways: (1) FactExtractionBlock(persist=True) writes them during a run (with a source trace); (2) the offline fastaiagent learn Trace Learning Loop mines them from past traces; (3) you insert them directly with MemoryStore.add(Fact(...)) (source = manual). The scope/scope_id values are whatever the producer assigned — they are not auto-detected from conversation.
Per-turn live memory (what a block recalled this turn, with scores) lives in the trace's memory.read spans above; the Memory page shows the durable facts that persist between runs.
Reproduce all three views with examples/memory_observability/ (companion.py seeds a trace, snapshot.py captures the UI).
Safety¶
Each block runs inside a try/except inside ComposableMemory. A failing block is logged and skipped — it cannot break the agent run. Individual block state survives across the failure.
Future work¶
Blocks are synchronous today: inside arun, a block's LLM, embedding and store calls run in line and hold the event loop while they do. Async block methods (aon_message, arender) are not shipped yet; when they are, they'll be additive — the sync API won't change.
Next Steps¶
- Agents — Core agent documentation
- KB Backends —
VectorStorebackends used byVectorBlock - Middleware — Transform agent messages and responses (complements memory)
- Tracing — Debug agent execution with traces
examples/customer-support-agent/—AgentMemorywired into a REPL so support sessions retain context across turns.