Connected HITL (observer)¶
Human-in-the-loop pauses and resolutions in the SDK are local mechanisms:
interrupt() suspends a run, your own app approves or rejects it, and
chain.resume() / agent.aresume() continues. See
Human-in-the-Loop for that core flow — it works fully
standalone, with no platform dependency.
When you fa.connect() to an Enterprise control plane, the SDK additionally
reports those pauses and resolutions to the plane so it can serve an
org-wide pending/paused status view and a compliance ledger across every
connected agent.
Observer model — the plane never approves¶
The plane is an observer only. It records that a run paused and how it was resolved; it is never the approval surface. Approval always happens in your own app (mobile / web / API / middleware), exactly as it does without a connection. Nothing about how you resolve a pause changes when connected.
What is reported¶
Reporting is metadata only — no interrupt payloads, no RunContext, no
user/business data ever leaves the process:
| Field | Pause | Resolution |
|---|---|---|
run_id |
execution id | execution id |
event_type |
paused |
resolved |
kind |
approval for a managed-policy pause, else interrupt |
same as its pause |
agent_id / chain_id |
which agent/chain | which agent/chain |
node |
node / turn:N/tool:name |
the resumed node |
reason |
the interrupt() reason label |
the original reason |
status |
— | approved / rejected |
resolver |
— | resume_value.metadata["resolver"] if set |
context |
— | policy pause only: {"pending_id": <id or null>}; otherwise none |
kind tells the two kinds of pause apart on the ledger: approval is a tool
call a managed approval policy paused (reason policy_approval_required, since
1.74.0 — earlier versions reported those as interrupt too); interrupt is an
interrupt() in your own code. status is only ever approved or rejected.
context.pending_id (since 1.76.0) names the pending run the plane registered
for a policy pause, so the plane closes exactly that pause instead of matching by
position in the run. It is null when registering the pending run failed, which
tells the plane there is nothing to close. It is the only thing context ever
holds.
The raw interrupt(reason, context) context dict is never sent — only the
short reason label. To attribute a resolution to a person, pass a resolver in
the resume metadata — the identity of whoever answered in your app. It is the
only source of that identity, and for a policy approval it is the heart of the
human-oversight evidence:
import fastaiagent as fa
from fastaiagent.chain.interrupt import Resume
fa.connect(api_key="fa-...", target="https://your-plane.example.com")
# ... agent/chain pauses on interrupt(); your app collects an approval ...
await chain.resume(
execution_id,
resume_value=Resume(approved=True, metadata={"resolver": "alice@acme.com"}),
)
# A `resolved` event (status=approved, resolver=alice@acme.com) is reported.
Local-first, never blocks¶
Reporting reuses the same durable outbox as trace export:
- On a pause/resolution the event is written to a local
hitl_eventstable (synced=0) — the durable source of truth. - A background drain POSTs un-acked events to
/public/v1/hitl/eventswith bounded retry (transient/5xx retried, 4xx terminal), marking themsynced=1only after a confirmed2xx. - Re-send is idempotent by a SDK-generated
event_id, so an outage that overlaps a partial send never double-counts.
The agent hot path is never blocked: the local write is fast and the network POST runs on a background thread. When not connected, reporting is a strict no-op — nothing is written and nothing is sent.
Bounded buffer. The re-send queue is capped (~100,000 un-acked events or
~30 days — gentler than traces, since HITL events are rare and
audit-significant). Beyond that the oldest are dropped from the re-send queue
but kept in local.db; the dropped count is logged.
Enablement¶
Connected HITL reporting is part of the Enterprise bundle, gated by the
connected_state_plane feature flag on your domain. If the domain is not
entitled, the ingest endpoint returns 403 — the SDK logs a warning, leaves the
events buffered (a terminal 4xx is not retried), and the agent runs unaffected.
Upgrade note: the local
hitl_eventstable is created by an automatic, additive migration (local schema v12). Existing projects are unaffected; only pauses/resolutions that occur while connected are reported.
In the console¶
The plane records every pause and resolution the SDK reports and serves an org-wide HITL status view across connected agents — paused / resolved, the node, the outcome, and the resolver (the approval surface stays your own app):

A runnable end-to-end example is in examples/85_connected_hitl.py.
Next steps¶
- Human-in-the-Loop — the core
interrupt()/resume()flow - Platform Connection —
fa.connect()and the other connected services - Managed governance (approvals) — policy-gated approvals