Skip to content

Human-in-the-Loop

There are two ways to gate a chain on a human decision.

Mode When to use API
Suspending HITL (interrupt()) Approvals that may take seconds, hours, or days. Process exits cleanly between pause and resume. interrupt(reason, context) inside any node + chain.resume(execution_id, resume_value=Resume(approved=True))
Blocking HITL (NodeType.hitl + handler) Inline CLI-style prompts that block in the calling process. Useful for tests and local scripts. chain.add_node("review", type=NodeType.hitl) + chain.execute(..., hitl_handler=fn)

The two modes are independent — pick the one that matches the workflow.

For end-to-end durability (suspending HITL plus crash recovery, the /approvals UI, Postgres in production, the @idempotent decorator that makes resumed nodes safe to re-run, and resume from Python / HTTP / CLI), see the dedicated Durability section.

Suspending HITL — interrupt()

Calling interrupt(reason, context) inside any node suspends the chain cleanly. The executor persists an interrupted checkpoint plus a row in pending_interrupts, and chain.execute(...) returns a ChainResult with status="paused". The Python process can exit. A separate chain.resume(...) call (in a different process, hours later) injects a Resume value and re-runs the suspended node — this time interrupt() returns the value instead of raising.

from fastaiagent import Chain, FunctionTool, Resume, interrupt
from fastaiagent.chain.node import NodeType

def approval(amount: int):
    if amount > 10_000:
        decision = interrupt(
            reason="manager_approval",
            context={"amount": amount, "policy": "high-value"},
        )
        return {"approved": decision.approved,
                "approver": decision.metadata.get("approver")}
    return {"approved": True, "auto": True}

chain = Chain("payments")
chain.add_node("approval",
               tool=FunctionTool(name="approval_tool", fn=approval),
               type=NodeType.tool,
               input_mapping={"amount": "{{state.amount}}"})
# ... wire up other nodes ...

# First run — suspends.
result = chain.execute({"amount": 50_000}, execution_id="payment-abc")
assert result.status == "paused"

# Sometime later (different process is fine):
result = await chain.resume(
    "payment-abc",
    resume_value=Resume(approved=True, metadata={"approver": "alice"}),
)
assert result.status == "completed"

Frozen-context invariant

The context dict you pass to interrupt() is JSON-serialized at suspend time and stored verbatim in the checkpoint and pending_interrupts row. On resume, the executor does not recompute it — the human approved a specific snapshot, and that snapshot is what the resumer sees.

This matters because the world may have changed between pause and resume: balances move, prices update, customers cancel. If your node needs the current values when it resumes, read them from chain state (which is also rehydrated on resume) instead of relying on the frozen context.

Three resume entry points

Chain.resume(...) is the Python API. The same internal path is reachable from two more places:

# CLI — useful for ad-hoc operator runs and ops scripts.
fastaiagent resume <execution-id> \
    --runner myapp.chains:my_chain \
    --value '{"approved": true, "metadata": {"approver": "alice"}}'

fastaiagent list-pending           # rich table of every paused workflow
fastaiagent inspect <execution-id> # checkpoint history for one execution
# HTTP — the /approvals UI (Phase 10) and any external system call this.
from fastaiagent.ui.server import build_app

# Pass the chains/agents the server needs to resume.
app = build_app(db_path=..., runners=[my_chain, my_agent, my_swarm])
# Same thing from the CLI. Without it, Approve/Reject returns 503 —
# the UI can show a pending interrupt but not resolve it.
fastaiagent ui --agent app.py:my_chain --agent app.py:my_agent
# POST /api/executions/{execution_id}/resume
curl -X POST http://127.0.0.1:7842/api/executions/refund-abc/resume \
    -H 'Content-Type: application/json' \
    -d '{"approved": true, "metadata": {"approver": "alice"}}'

All three converge on the same atomic-claim semantics: concurrent resumers see AlreadyResumed (CLI exit code 2; HTTP 409 Conflict).

Side effects before interrupt()

Because resume re-runs the suspended node from the top, anything you do before the interrupt() call runs twice. If your node charges a card, sends an email, or makes any other side-effectful call before suspending, wrap it with @idempotent — that's exactly what the decorator exists for.

Sibling tool calls in the same turn

When an agent's model asks for several tools in one turn and the first of them calls interrupt(), resume answers that one and continues at the next turn. The siblings are not re-dispatched: their side effects never happen, and each is answered with a note telling the model the tool did not run. Resume re-enters a run at a checkpoint, not in the middle of a turn, so there is nowhere to re-enter that would run them in order — and running them after the fact would fire side effects for calls the model never saw a first result for.

If a particular tool must survive a pause, put the interrupt() in that tool, or drive the calls across separate turns.

Atomic resume claim

chain.resume(execution_id, resume_value=…) atomically deletes the pending_interrupts row before re-running the suspended node. If the row is already gone (another resumer beat us, or the chain was never suspended), it raises AlreadyResumed. This is the safety net against double-clicking an "Approve" button.

from fastaiagent import AlreadyResumed

try:
    await chain.resume(execution_id, resume_value=Resume(approved=True))
except AlreadyResumed:
    print("This approval was already processed.")

Blocking HITL — inline handler

Use this for CLI tooling, tests, or scripts where the chain runs to completion in a single process and a human is sitting at the terminal.

Basic Usage

from fastaiagent import Agent, Chain, LLMClient
from fastaiagent.chain import NodeType

chain = Chain("approval-pipeline")
chain.add_node("draft", agent=drafter_agent)
chain.add_node("review", type=NodeType.hitl)
chain.add_node("send", agent=sender_agent)
chain.connect("draft", "review")
chain.connect("review", "send")

Custom Approval Handler

Provide a handler function that receives the node, context, and state, and returns True (approve) or False (reject):

def approval_handler(node, context, state):
    draft = context["node_results"]["draft"]
    print(f"Review this draft: {draft}")
    return input("Approve? (y/n): ").lower() == "y"

result = chain.execute({"message": "Write a response"}, hitl_handler=approval_handler)

The handler has access to: - node — the HITL node definition - context — includes node_results from all previously executed nodes - state — the current chain state

The handler may be async def; it is awaited. Its answer is read as a yes or a no — anything that is not truthy, including a handler that forgot to return, is a no.

Rejection stops the chain

When the handler says no, the chain stops at the gate. Nothing after it runs, and the result says so:

result = chain.execute({"message": "..."}, hitl_handler=my_handler)
if result.status == "rejected":
    print(result.node_results["review"])   # {"approved": False}
    print(result.output)                   # None — the chain did not finish its work

A rejected run has ended: resuming it raises AlreadyResumed, like a completed one.

Changed in 1.77.0

Until 1.77.0 a rejection recorded approved=False and the chain carried on — send ran unless the author had wired a condition node after the gate to catch it, and getting that wiring slightly wrong sent the rejected draft anyway. That workaround is no longer needed. If you had one, it now never sees a rejection (the chain stops first), so anything its "abort" branch did — a notification, a log line — belongs after checking result.status == "rejected" instead.

You can inspect the approval decision after execution:

result = chain.execute({"message": "..."}, hitl_handler=my_handler)
review_result = result.node_results.get("review", {})
print(review_result.get("approved"))   # True or False
print(review_result.get("message"))    # "Auto-approved (auto_approve=True)" if opted in

An unconfigured gate refuses

A NodeType.hitl node with no hitl_handler raises ChainError and fails the run. Changed in 1.67.0.

Until then it approved itself and returned {"approved": True, "message": "Auto-approved (no HITL handler)"}, and the run reported status="completed". That is a control which could not run reporting a clean pass — in the one node type whose entire job is to stop things. Forgetting to pass hitl_handler to a call that did pass it in development was enough.

chain.add_node("review", type=NodeType.hitl)
chain.execute({"message": "..."})          # ChainError: Approval gate 'review' has no handler
chain.execute({"message": "..."}, hitl_handler=my_handler)   # fine

A pause needs somewhere to be saved

An interrupt() in a chain built with checkpoint_enabled=False has nowhere to store the pause, so it is not reported as status="paused" — a resume of it could never work. Since 1.77.0 the InterruptSignal rises to the caller instead (or to whatever encloses the chain and can hold it). Keep checkpointing on for any chain that pauses.

A pause does not expire

A pause stays open until your application resumes it: the SDK sets no deadline and never expires or decides one on its own. That is deliberate — the decision is yours, and a timeout that silently approved or rejected would make it for you. If a pause should not wait forever, resume it from your side with the answer your policy calls for. A connected plane may flag a pause that outlives its approval policy's timeout, but it never decides it either.

Auto-approve, when you ask for it (testing)

The convenience survives; it just has to be requested on the node, where it is visible in the chain's definition and in its serialized form:

chain.add_node("review", type=NodeType.hitl, auto_approve=True)
result = chain.execute({"message": "Write a response"})   # approves, says so

Or pass a lambda for quick testing, which needs no opt-in because it is a handler:

result = chain.execute(
    {"message": "Write a response"},
    hitl_handler=lambda n, c, s: True,  # Always approve
)

Complete Example

A support pipeline with drafting, review, and sending:

from fastaiagent import Agent, Chain, LLMClient
from fastaiagent.chain import NodeType

llm = LLMClient(provider="openai", model="gpt-4.1")

chain = Chain("support-pipeline")
chain.add_node("draft", agent=Agent(
    name="drafter", system_prompt="Draft a helpful response.", llm=llm))
chain.add_node("review", type=NodeType.hitl)
chain.add_node("send", agent=Agent(
    name="sender", system_prompt="Finalize and send the response.", llm=llm))
chain.connect("draft", "review")
chain.connect("review", "send")

def review_handler(node, context, state):
    draft = context["node_results"]["draft"]
    print(f"\n--- Draft for review ---\n{draft}\n---")
    decision = input("Approve? (y/n): ").strip().lower()
    return decision == "y"

result = chain.execute(
    {"message": "My order hasn't arrived"},
    hitl_handler=review_handler,
)
print(result.output)

Next Steps