Skip to content

Replay Fidelity Guarantees

This page documents what Agent Replay captures, what it can faithfully reproduce, and where the known fidelity gaps are. Each guarantee is backed by a test in tests/test_replay_determinism.py.

Determinism modes

ForkedReplay.with_determinism(mode) controls how the LLM is invoked during a rerun:

Mode LLM HTTP call Output
"live" (default) Yes — uses the captured temperature/seed. Nondeterministic for non-zero temperature. Matches the original behavior.
"recorded" No — replays the captured response. Byte-identical to the original. Use for regression tests.
"deterministic" Yes — but with temperature=0 and seed=42 forced. Semantically stable across runs where the provider supports seed.
from fastaiagent.trace.replay import Replay

replay = Replay.load(trace_id)
forked = (
    replay
    .fork_at(step=0)
    .modify_prompt("Be more concise.")
    .with_determinism("recorded")  # ← skips the LLM call entirely
)
result = forked.rerun()

What "recorded" mode captures

The recorded LLMResponse is reconstructed from these GenAI semantic-convention attributes on the captured llm.* span:

  • gen_ai.response.content — main text output
  • gen_ai.response.finish_reason — stop, length, etc.
  • gen_ai.response.tool_calls — tool invocations the original LLM requested (JSON-decoded into ToolCall instances)

If FASTAIAGENT_TRACE_PAYLOADS=0 was set on the original run, none of these attributes are present and recorded mode raises ReplayError. The fix is to enable payloads on the runs you intend to replay later.

A turn that only calls tools is replayed too (1.84.0). OpenAI sends content: null on such a turn, so its span has gen_ai.response.tool_calls and no gen_ai.response.content. Before 1.84.0 those spans were skipped. A recorded rerun of any tool-using agent jumped straight to the captured final answer, never ran a tool, and still matched the original output. Now the tool-call turn is served in order, so the rerun really executes the agent's tools. Use with_tool_override(...) to stub any tool with side effects.

Decisions API calls (1.84.0)

Calls to LLMClient.adecide are recorded on their own llm.<provider>.decisions.<model> spans (identified by gen_ai.operation.name="decisions"), and they replay from their own queue. A decision is never handed to a chat turn, or a chat turn to a decision. Under recorded, each adecide gets the next captured result, tagged replay.mode="recorded". A decision the original run didn't make is a miss, handled by on_miss. That covers more decisions than were captured, or a captured one that answered different questions.

A recorded rerun: chat turns and the Decisions call all served from the capture

Multi-turn replay (v1.14.1+)

For a multi-turn tool-loop trace with N captured llm.* spans, recorded mode installs the full ordered sequence of responses (sorted by start_time) and pops the front entry for each LLMClient.acomplete call during rerun. So:

  • Turn 1 of the rerun gets the trace's first captured response.
  • Turn 2 gets the second, etc.
  • If the rerun makes more LLM calls than the original captured (e.g. a tool override that triggers an extra reasoning turn), the queue drains and subsequent calls fall through to a live provider call — the agent doesn't deadlock.

Pre-v1.14.1 the same first response was returned for every LLM call in the rerun, which made tool-loop replays nonsensical (every turn parroted turn 1).

Controlling the drained-queue fall-through (v1.48.0+)

That live fall-through is billed and nondeterministic, and before v1.48.0 it happened silently. It is now warned loudly, and you can make it fatal — recommended whenever the rerun is a regression test:

forked.with_determinism("recorded", on_miss="error")   # raises ReplayError
forked.with_determinism("recorded")                    # default: on_miss="live"

on_miss="error" raises before any provider call, so a recorded suite can never quietly spend tokens or return drifted output.

Known fidelity gaps

These are documented limitations to set expectations. Each is tracked in the top-3-recommendation as a deferred follow-up.

Streaming chunks aren't recorded granularly

Streaming responses (agent.astream(...)) yield individual StreamEvent chunks, but only the final accumulated content lands in gen_ai.response.content. recorded mode reconstructs the LLMResponse from that final text — streaming cadence is lost.

For most regression-test scenarios this is fine: the final answer is what you're testing. If you need per-chunk fidelity, file an issue — granular streaming capture is on the deferred-follow-up list.

Tool execution is rerun, not recorded

recorded mode skips the LLM call but does not skip tool calls. If the LLM's recorded response includes a tool invocation, that tool will execute live during rerun. Two ways to handle this:

  • with_tool_override(name, stub) — replace a single tool with a deterministic stand-in. Other tools keep their reconstructed implementations. Multiple calls compose.
  • with_tools([...]) — full replacement. All tools come from the list you pass; nothing is reconstructed from the trace.
forked.with_tool_override("search_kb", deterministic_stub)
forked.with_tool_override("create_ticket", another_stub)

Pauses and governance during a rerun

The rerun agent is rebuilt from the trace without a checkpointer and without its platform agent_id. Two consequences:

  • A rerun cannot pause. If a tool calls interrupt(), arerun() raises ReplayError naming what it paused on; override that tool with with_tool_override(...) to replay past it. (Before 1.78.0 the bare InterruptSignal escaped arerun().)
  • A rerun is not governed. Managed approval policies never gate it. Rebuilt tools have no function unless you supply one, so a real side-effecting tool only runs in a rerun when you pass it in through with_tool_override / with_tools — an explicit, local choice.

Marking tools with replay_class

Rather than override tools by hand on every rerun, declare each tool's replay-safety class once at definition time. The central Replay engine reads it (off the fastaiagent.tool.replay_class span attribute) to decide, per call, whether to re-execute or inject the recorded output:

replay_class Replay behavior
"read_only" May be re-executed (no side effects — a GET, a pure lookup).
"idempotent" Recorded output is injected (re-running is safe but unnecessary).
"side_effecting" Recorded output is injected; never re-executed.
from fastaiagent.tool import tool

@tool(name="search_kb", replay_class="read_only")
def search_kb(query: str) -> str:
    ...

@tool(name="create_ticket")           # unmarked → "side_effecting" (safe default)
def create_ticket(summary: str) -> str:
    ...

The default is "side_effecting", and a value is never auto-inferred — a GET RESTTool is not automatically read_only. See Tools → Replay safety.

Provider-specific determinism support

Provider temperature=0 Seed support
OpenAI / Azure OpenAI ✅ ✅ via seed=
Anthropic ✅ ✅ via seed=
Gemini (native wire) ✅ ✅ via seed=
Bedrock ✅ ❌ — Bedrock has no seed field
Ollama / vLLM / local ✅ ✅ via seed=
Provider presets (groq, openrouter, mistral, …) ✅ Varies — most pass through seed=

deterministic mode degrades gracefully on Bedrock — it forces temperature=0 and skips the seed. Output is still much more reproducible than live but not byte-identical across runs.

Divergence detection

forked.compare(result) now computes the first divergence by walking the original and rerun step lists in parallel. Span names and span output attributes are compared step-for-step; the first mismatch wins.

This replaces the v1.13 behavior of hardcoding diverged_at to the fork point — which was misleading because the actual divergence often happens later than where the user asked to fork.

comp = forked.compare(rerun_result)
assert comp.compare_status == "ok"
if comp.diverged_at is None:
    print("rerun matched the original step-for-step")
else:
    print(f"diverged at step {comp.diverged_at}: "
          f"{comp.original_steps[comp.diverged_at].span_name}")

When the rerun trace can't be loaded (e.g., Agent.arun failed), compare_status is set to "rerun_failed" and diverged_at stays None — so UIs can distinguish "matched everything" from "couldn't tell".

Saving a rerun as a regression test

rerun.save_as_test(...) appends a JSONL line in the format evaluate() reads natively:

rerun.save_as_test(
    "regression_tests.jsonl",
    input="What is our refund policy?",
    expected_output=str(rerun.new_output),
    source_trace_id=original_trace_id,
)

For the full failure → fix → regression-test loop, see examples/62_replay_to_regression.py.