Concepts & Mental Model¶
This page explains how durability actually works under the hood — the checkpoint record, exactly when it's written, how resume decides where to re-enter, and the atomic operations that make interrupt/resume and idempotency safe. It's an internals explainer, not a quickstart; for setup see the Durability reference and Quickstart.
The problem¶
An agent or chain is a sequence of steps, some of which are expensive (LLM calls), side-effecting (charge a card), or slow (wait days for a human). If the process crashes at step 4 of 6, you don't want to redo steps 1–3. If a step pauses for human approval, you don't want to keep a process alive for three days. And when execution does re-enter a step, you must not fire its side effects twice.
Durability solves all three by writing the run's progress to a checkpoint store as it goes, so any process with access to that store can resume exactly where the last one left off.
The checkpoint record¶
Everything rests on one row. A Checkpoint (fastaiagent/chain/checkpoint.py)
captures the state of a run at one boundary:
| Field | What it holds |
|---|---|
execution_id |
The run identifier — immutable across resume; the key everything is looked up by |
node_id |
Where this checkpoint is — the re-entry address (see the scheme below) |
node_index |
Position in topological/turn order |
step_type |
What kind of boundary this is: llm_call, tool_call, hitl_pause, node, handoff, fork_origin, or run_end |
status |
"completed", "interrupted", or "failed" — this drives resume |
state_snapshot |
The full run state at this point, JSON-frozen (chain state, or serialized messages + turn for an agent) |
node_input / node_output |
The step's inputs (e.g. saved tool args) and result |
iteration_counters |
Cycle counts, so a resumed chain loop picks up mid-cycle |
interrupt_reason / interrupt_context |
For a paused run — the frozen snapshot the human approves |
agent_path |
Hierarchical location in multi-agent topologies (see composition) |
The node_id is the re-entry address, and its scheme tells you what kind of
boundary it is:
- Chain node — the node's own id (e.g.
"review"). - Agent turn —
"turn:N"(checkpoint written before the Nth LLM call). - Agent tool —
"turn:N/tool:<name>"(written before the tool runs, with the tool's args innode_input). - Run end —
"run_end", withstep_type="run_end". Not an address at all; a tombstone. See below.
When checkpoints are written¶
Checkpoints are written at boundaries chosen so resume never loses or repeats committed work:
- Chain: after each node completes → a
status="completed"checkpoint with the post-nodestate_snapshot(chain/executor.py). - Agent, per turn:
_put_turn_checkpointruns before each LLM call (turn:N) — the resume point for a crash mid-inference. -
Agent, per tool:
_put_tool_checkpointruns before each tool dispatch (turn:N/tool:X), saving the exact args — so resume re-invokes the tool with the same input. -
Run end: one terminal row per run,
step_type="run_end", withstatus="completed"on success or"failed"when an exception escaped — including an exception from after the tool loop (the re-ask, output guardrails, memory). It is written by the outermost runner only — a Swarm's child agents and a Supervisor's workers share their parent'sexecution_id, and a marker per hop would claim the run ended at every handoff. A paused run gets none: a pause is not an ending.Since 1.68.0 a fork is a run by this rule too.
Chain.aforkandAgent.aforkwrite the same terminal row under the branch's ownexecution_id—failedwhen the branch raises,completedwhen it finishes, and none when it pauses. Before 1.68.0 a forked branch wrote no marker at all in either direction, so a branch that died was byte-identical to one that merely stopped, and a branch that finished left nothing to stoparesumere-executing it.
Why the run-end row exists: without it a run that finished and a run that died
the instant after its last step are byte-identical — both leave a completed
checkpoint as the newest row. Nothing could tell them apart, so aresume on an
already-finished run silently re-executed it, re-calling the model and
re-firing every side effect not wrapped in @idempotent. It now raises
AlreadyResumed instead. A failed marker is stepped over rather than refused —
crash recovery is the whole point, so a run that raised stays resumable.
Fork reads the same rows and answers differently, which is the one place the
tombstone's meaning splits. resume re-enters the same run, so a finished one
must be refused. fork writes a new execution_id and leaves the original
untouched, so a finished run is the most ordinary thing to branch. Note which
run each question is about: once a branch has its own marker (1.68.0), resuming
the branch follows the ordinary rule — a completed branch raises
AlreadyResumed, a failed one is stepped over and stays resumable. Both must skip
the marker — it addresses no node — and they do so through two helpers,
latest_resumable and latest_forkable. Until 1.67.0 the fork paths skipped
nothing: they called get_last, picked the tombstone, and refused with
"Cannot fork … from node 'run_end': it is the final node", which is false twice
over — run_end is not a node, and the run it was refusing to branch was often a
failed one that had branched fine before the marker existed.
execution_id is minted once at the start of a run (or supplied by you) and
placed in a ContextVar so every node, tool, and @idempotent function in that
run reads the same id.
Verified against a live run
A chain node that called interrupt() wrote exactly one checkpoint —
node_id='review' status='interrupted' interrupt_reason='manager_approval'
— plus a pending_interrupts row, and the result came back
status="paused" with pending_interrupt={reason, context, node_id,
agent_path}. That is the whole suspended-run footprint on disk.
How resume decides where to re-enter¶
Resume reads the latest checkpoint's status and branches. This is the
contract (Chain.aresume / Agent.aresume):
| Latest status | You pass | What happens |
|---|---|---|
interrupted |
resume_value=Resume(...) |
Atomically claim the pending row, then re-run the interrupted node from the top with the resume value in scope — so interrupt() returns it instead of raising |
interrupted |
nothing | Raises ChainResumeError — an interrupted run needs a Resume(...) |
completed / failed |
nothing | Restart at the next node after the last committed one |
The subtle, important part is the interrupted case: the paused node is
re-executed from the beginning. Everything before the interrupt() call
runs again — which is exactly why side effects need @idempotent (below). The
resume value is injected via a ContextVar set only for that first
re-executed node, then cleared, so downstream nodes run normally.
For an agent, the same idea specializes by node_id: a turn:N crash re-issues
the LLM call with saved history; a turn:N/tool:X crash re-invokes the tool
with saved args (no LLM re-call); a tool interrupt() re-invokes the tool so
its interrupt() returns the Resume.
A turn with several tool calls¶
The pre-tool checkpoint is written before dispatch, so when a model asks for three tools in one turn and the first pauses, the saved history holds the assistant message declaring all three and no results at all.
Resume answers the call it suspended on and then continues at the next turn. The siblings are not re-dispatched — resume re-enters a run at a checkpoint, never in the middle of a turn, and firing them here would run side effects the model never saw a first result for. They are instead answered with a note saying the tool did not run, so:
- the history stays acceptable to OpenAI and Anthropic (before 1.67.0 it was not, and the 400 arrived on the first request after the resume — from a shape that had been persisted, so it survived the process that made it);
- the model is told plainly that the work did not happen, and can ask for it again on the next turn if it still needs it.
If a tool must run even across a pause, make it the call the pause lands on, or drive the calls across separate turns.
A failure after the tool loop¶
A run can also die past the loop — in the structured-output re-ask, an output
guardrail, or a memory write. Since 1.67.0 that writes a failed run-end marker
like any other crash (before, it wrote none at all, and the run was
indistinguishable from one that simply finished), plus a turn-boundary row one
past the loop's last turn. Resume then re-enters after the loop and re-issues
the model, rather than stepping back into the last pre-tool checkpoint and
re-invoking a tool that had already run.
Interrupt, suspend, and the atomic claim¶
interrupt(reason, context) (fastaiagent/chain/interrupt.py) is the whole
suspend mechanism, and it's tiny:
def interrupt(reason, context):
v = _resume_value.get()
if v is not None:
return v # resuming: return the human's decision
raise InterruptSignal(reason, context) # first pass: suspend
First pass, there's no resume value, so it raises InterruptSignal. The
executor catches it and, in one transaction (record_interrupt), writes
both the interrupted checkpoint and a pending_interrupts row — so the
approvals UI never sees a half-suspended run — then returns status="paused".
The process can now exit.
Frozen context: the context dict is JSON-serialized at suspend time and
never recomputed. The human approves a specific snapshot, not whatever the world
looks like at resume time.
Double-resume safety: resume first claims the pending row by deleting it,
and the claim is atomic. On SQLite it's a BEGIN; SELECT; DELETE; COMMIT under
a lock; on Postgres it's a single DELETE ... RETURNING. Exactly one caller
gets the row back; everyone else gets None and raises AlreadyResumed. This
is what makes "resume" safe under concurrent resumers and against a
resume-after-completion.
Verified against a live run
Resuming with Resume(approved=True) claimed the row, re-entered review,
and interrupt() returned the decision → run completed. A second
resume for the same execution_id raised AlreadyResumed.
Idempotency — absorbing the re-execution¶
Because a resumed node runs from the top, any side effect before the resume
point fires again. That's the footgun. @idempotent
(fastaiagent/chain/idempotent.py) absorbs it:
On first execution within a run it runs the body and caches the (JSON-serialized)
result under (execution_id, key). The default key is
sha256(qualname + args + kwargs). On any later call in the same
execution_id with the same key, it returns the cached value and never runs
the body again. In a different execution, or outside any chain run (e.g. a
unit test), it's a cache miss and runs normally.
Verified against a live run
With charge_card wrapped in @idempotent, running → suspending →
resuming a high-value charge fired the side effect exactly once
(calls == 1), even though the review node re-executed on resume.
The backends¶
A checkpointer is anything implementing the Checkpointer protocol
(put, get_last, list, record_interrupt,
delete_pending_interrupt_atomic, get_idempotent/put_idempotent, prune).
Two ship in-box, over three tables (checkpoints, pending_interrupts,
idempotency_cache):
SQLiteCheckpointer(default) — single-file, great for local/dev and single-process. Atomicity comes from an in-process lock + explicit transactions. Cross-process coordination relies on SQLite file locking (not for distributed use).PostgresCheckpointer— for multi-process/production. The atomic claim is aDELETE ... RETURNINGand correctness rests on Postgres MVCC, so many workers can race to resume and exactly one wins.
prune(older_than) deletes old completed/failed checkpoints and idempotency
rows but never interrupted ones — pruning a suspended run would orphan a
live human-in-the-loop.
Composition across agents, chains, swarms¶
All checkpoints for a run share one execution_id; the agent_path column says
which agent/worker wrote each one. It's a ContextVar that's extended, not
overwritten, as topologies nest:
| Topology | agent_path shape |
|---|---|
| Chain | (none; node_id is the rendezvous point) |
| Agent | agent:<name> → agent:<name>/tool:<tool> |
| Swarm | swarm:<s>/agent:<a>/tool:<t> |
| Supervisor | supervisor:<s>/worker:<r>/tool:<t> |
So a Supervisor and its workers checkpoint into the same run, and resume can
scope to one agent's subtree by matching the agent_path prefix.
afork vs aresume¶
aresume(execution_id, ...)continues the same run from its last checkpoint (crash recovery, or returning aResume).afork(execution_id, checkpoint_id=..., ...)branches a newexecution_idfrom a saved checkpoint (linked viaparent_checkpoint_id), leaving the original intact — for counterfactuals and what-if replays. (This is the checkpoint-based cousin of trace-based Replay, which chains can't use.)
Next steps¶
- Durability reference · Quickstart
- Side effects & idempotency — patterns for
@idempotent - Multi-agent —
agent_pathacross Swarm/Supervisor - Checkpointers — SQLite vs Postgres, custom backends
- Chains — HITL · Agents — durability