AutoLLM¶
AutoLLM (fastaiagent.optimize) is eval-driven prompt optimization. Where
harden() recommends prompt fixes, AutoLLM closes the loop: it proposes a
change, applies it to a fresh agent, re-evaluates, keeps the best, and repeats —
until the score stops improving or a budget runs out. A held-out split guards the
winner against overfitting.
It tunes the system prompt by default, and can also tune few-shot
examples and which learned-memory facts to inject when you opt in — greedy
coordinate ascent, cycling the active levers one per round. The SDK's answer to
LangSmith's Promptim / DSPy's BootstrapFewShot + metaprompt optimizers, built
on the evaluate() you already use.
This is the OSS on-ramp: standard prompt optimization grounded in your own eval
data, end to end in one SDK. A runnable, real-LLM walkthrough lives in
examples/autollm/.
Scope
AutoLLM tunes the system prompt (default) plus, opt-in, few-shot examples and
which learned facts to inject — all on the cold-eval path. Runs are
persisted and viewable in the Local UI under AutoLLM. Replay-grounded
scoring (forking a production trace and rerunning from real operational state)
is the Enterprise complete-loop capability — the score_candidate seam is its
drop-in point.
Optimizing on your traces = agent-attributable cases only
When you build the eval set from traces
(curation), AutoLLM optimizes only on agent-quality failures,
not infrastructure failures: a run that infra-errored (endpoint 500, timeout)
and produced no usable output is dropped, not curated as a gold target, and a
case that errors during scoring is never shown to the prompt proposer as a
failure to fix. So the optimizer never chases a fault the agent can't fix — but
the errored case still counts against the candidate's score (see
When a case errors). Runnable walkthrough:
examples/80_curate_from_traces.py.
Quickstart¶
import fastaiagent as fa
agent = fa.Agent(name="capitals", system_prompt="You answer questions.", llm=fa.LLMClient())
report = fa.optimize(
agent,
"cases.jsonl", # Dataset | path | list[dict] with input/expected_output
scorers=["exact_match"],
config=fa.OptimizeConfig(max_iterations=5, patience=2),
)
print(report.summary())
better_agent = report.apply_to(agent) # a fresh agent with the winning prompt
optimize() is the sync wrapper; aoptimize() is the async implementation (it's
a minutes-to-hours operation — prefer async in apps).
How the loop works¶
split (seeded) → train / dev / holdout
baseline scored on dev (already at target_score → stop here)
repeat (cycling active levers: instructions → fewshot → memory):
propose N candidate variants of the active lever, on top of the current best
score each on dev
keep the best if it beats the current best by ≥ min_delta (else → patience)
stop on: patience | max_iterations | target_score | budget | proposer_failed
holdout guard: re-score the winner on the held-out split; revert to baseline
if it regressed beyond holdout_regression_tol
"Optimized" means hill-climbed until no improvement or budget exhausted — the same operational definition Promptim and DSPy use. The holdout guard, not the search, is what makes the result trustworthy rather than overfit: the holdout split never influences selection, so the reported lift is on data no candidate was tuned against. By construction the winner is never worse than baseline.
Algorithm
AutoLLM uses greedy coordinate ascent (Promptim-style keep/revert, one
lever per round) with a metaprompt / reflective proposer — the optimizer
reads the dev failures and writes a revised prompt. This is the same algorithm
family as LangSmith Promptim and DSPy (BootstrapFewShot + metaprompt
optimization). The joint-Bayesian-search variant — DSPy's MIPRO — searches
instructions and demos together; it's a documented upgrade path
(strategy="mipro") rather than the default, since coordinate ascent gives a
single-lever cause for every accepted step and avoids MIPRO's cost multiplier.
Each failing case shown to the proposer includes its expected output and the
scorer's reason (e.g. "got 1120, expected 1120000"), not just the input and
the wrong answer. This is what lets AutoLLM optimize extraction and
structured-output tasks, where the fix is an output convention (scale, sign,
formatting) the proposer can only infer by seeing what correct looks like — for
example recovering "values are in thousands → multiply by 1,000; parentheses are
negative; answer with the number alone" when pulling figures from financial
tables. (Added in 1.38.0; classification never needed it because the label space
is small enough to guess.)
The levers¶
instructions— rewrites the system prompt. The proposer reuses the failure analysis behindharden()but lives infastaiagent.optimize;harden()and the rest of theevalAPI are unchanged. The proposer is shown up to 40 failing train cases per round.optimize()raisesValueErrorfor an agent with a callable (dynamic)system_promptwhile this lever is active — leave"instructions"out ofleversfor such an agent.fewshot— bootstraps few-shot examples (DSPyBootstrapFewShot): gold(input, expected_output)pairs from the train split (pluscurate_from_traces(filter="favorites")), filling any gap by running the agent and metric-filtering its passing outputs. Demos are injected via aFewShotBlock. No demo ever carries a dev or holdout input — a favorite trace whose input is a scored case is skipped, so an eval set curated from your favorites can't hand the agent the answers it is scored on.memory— tunes which subset of the agent's learned facts to inject, via a confidence/recency ablation. It reads the facts where the agent's memory reads them: aMemory'slocationandproject_id(its global tier), or aPersistentFactBlock's own store, scope andproject_id; an agent with neither useslocal.dbat("agent", <agent name>). Pure selection — it never creates, edits, or deletes facts, so the audit chain is untouched. Injected through aPersistentFactBlockbacked by an allowlist over that same store. With no facts there it's skipped (recorded distinctly from a reject).fastaiagent learnwrites tolocal.db, so for an agent whose memory lives in Postgres or Redis, put the facts in that store.
The default is prompt-only (levers=("instructions",)) — the cheapest entry
point (few-shot adds a bootstrap pass; memory needs fastaiagent learn to have
run). Opt into more with e.g. levers=("instructions", "fewshot", "memory").
Memory-bearing agents¶
Each candidate evaluation gets an isolated copy of the agent's memory
(block.isolated_copy(): shares external handles like the llm, resets
in-process state) so one candidate's turns never bleed into another's. Agents
with StaticBlock / PersistentFactBlock / PlaneFactBlock / SummaryBlock /
FactExtractionBlock(persist=False) are supported. VectorBlock and
FactExtractionBlock(persist=True) are excluded — they write to an external
store during a run, so sharing them bleeds candidates; optimize() refuses
unless you pass allow_writable_memory=True (accepting the bleed).
Agents built on Memory¶
An agent whose memory is a Memory optimizes the same way, and the agent you get back still has a Memory — with the same window, location, agent_id and per-user routing:
from dataclasses import dataclass
import fastaiagent as fa
agent = fa.Agent(
name="support",
system_prompt="You answer billing questions.",
llm=fa.LLMClient(),
memory=fa.Memory(agent_id="support", user_id=lambda ctx: ctx.state.user_id, window=20),
)
report = fa.optimize(
agent, "cases.jsonl", scorers=["exact_match"],
config=fa.OptimizeConfig(levers=("instructions", "fewshot", "memory")),
)
better = report.apply_to(agent) # better.memory is a Memory, one window per user
@dataclass
class Session:
user_id: str
better.run("Why was I charged twice?", context=fa.RunContext(state=Session(user_id="alice")))
A runnable version, with the memory lever over facts in the agent's own store: examples/101_optimize_memory_agent.py.
What carries over and what is fresh:
- Configuration carries over: the store,
window,agent_id,max_users, and the per-user resolver — every user still gets their own window. - Windows start empty: each candidate, and the returned agent, starts with no conversations in memory. Durable facts are in the store and are read as usual.
- The levers reach every user: the few-shot demos and the selected facts are injected for each user, and for callers with no user.
- The memory lever works on global facts: it selects among the agent's global facts (
agent_id=, or the agent's name when unset) in theMemory's own store andproject_id— never among one user's facts.
Memory that writes during a run can't be isolated per candidate, so, as with the blocks above, optimize() refuses it unless you pass allow_writable_memory=True:
| Keyword | During optimize |
|---|---|
learn= |
refused — it writes learned facts to the store on every turn |
recall=<VectorStore> |
refused — every candidate would write to the same store |
recall="auto" |
allowed — each candidate builds its own in-process index |
summarize=, semantic= |
allowed |
Configuration¶
fa.OptimizeConfig(
max_iterations=8, # hard cap on rounds
patience=3, # stop after N non-improving rounds
target_score=None, # stop early once dev reaches this
candidates_per_iteration=3, # proposals per round
min_delta=0.01, # improvement smaller than this = "no improvement"
splits=(0.5, 0.25, 0.25), # train / dev / holdout
holdout_regression_tol=0.0, # revert if holdout drops more than this
seed=0, # deterministic split
primary_metric=None, # scorer name to select on (default: overall pass-rate)
max_eval_runs=None, # hard cap on evaluation passes (see Cost)
max_judge_calls=None, # hard cap on model-backed scorer calls (see Cost)
selection_judge=None, # an LLM judge used *inside* the loop
audit_judge=None, # an LLM judge used *only* on the holdout guard
levers=("instructions",), # default: prompt only — add "fewshot" and/or "memory"
allow_writable_memory=False, # opt in to memory that writes during a run (bleed risk)
)
The two-judge guard¶
For agents graded by a deterministic scorer, set primary_metric and you're
done. For reference-free agents (research, summarization, KYC narratives)
selection is an LLM judge — and optimizing against the same judge you report is
reward-hacking waiting to happen. Pass distinct judges:
from fastaiagent.eval import GEval
cfg = fa.OptimizeConfig(
selection_judge=GEval(criteria="answer quality"), # drives accept/reject
audit_judge=GEval(criteria="answer quality", name="audit", # different prompt/model
evaluation_steps=[...]),
)
If you leave audit_judge unset, it falls back to the selection judge with a
warning — fine for a first pass, not for a number you'll quote. Judges are
ordinary Scorers; they're composed into the scorers list (and deduped, so a
judge you already pass in scorers isn't billed twice).
Because judges are deduped by name, an audit_judge must not share its name
with a different scorer in scorers — results are keyed by name, so the holdout
would be scored by that other scorer instead. optimize() raises ValueError up
front when they clash: LLMJudge defaults to "llm_judge" and GEval to
"g_eval", so give the audit judge its own name=, as above. Passing the same
judge object in scorers and as audit_judge makes it drive selection as well, and
optimize() warns.
Reading the report¶
OptimizationReport mirrors HardeningReport (.summary(), .to_dict()) and
adds the score trajectory and an applyable winner:
Optimization — capitals (stopped: patience)
============================================================
baseline dev=0.600
iter 1 [instructions] dev=0.800 (+0.200) ACCEPT — answer with only the place name
iter 2 [instructions] dev=0.800 (+0.200) reject
------------------------------------------------------------
best dev=0.800
holdout best=0.750 (baseline=0.600, Δ+0.150) → winner kept
report.best_candidate.system_prompt— the winning prompt.report.apply_to(agent)— a copy of your agent with the winning levers applied: same class, tools, guardrails, middleware and agent path. The original is never mutated. A changed prompt dropsprompt_slug, since the registry prompt it names is no longer what the agent runs.report.trajectory— every candidate scored, with lever attribution and the number of dev cases thaterrored.report.improved— did the winner beat baseline and survive the holdout guard?report.stopped_reason—patience,max_iterations,target_score,budget,proposer_failedorno_active_levers, with+revertedappended when the holdout guard reverted the winner.report.proposer_errors— each time the prompt proposer could not run or its reply could not be read.
When a case errors¶
A case that raises instead of answering — a guardrail block, MaxIterationsError,
a provider error — counts as a failure in the candidate's dev and holdout
scores (score 0 on every metric). evaluate() leaves such a case out of its own
pass rate, but selecting on that would let a candidate that crashes on its hard
cases outscore one that answers them. The summary shows the count:
An errored case is still never shown to the prompt proposer as a failure to fix. A flaky provider therefore costs a candidate points rather than handing it a win: use a model client with retries for long runs.
When the proposer fails¶
If the prompt proposer can't run — an unknown model, an auth error, a reply that
isn't the requested JSON — the round is recorded as a skipped step, the error is
logged as a warning and kept in report.proposer_errors, and a run that ends
because of it stops with proposer_failed, not patience:
Optimization — capitals (stopped: proposer_failed)
============================================================
baseline dev=0.000
iter 1 [instructions] SKIPPED — proposer failed: LLMProviderError: OpenAI API error 404 …
iter 2 [instructions] SKIPPED — proposer failed: LLMProviderError: OpenAI API error 404 …
------------------------------------------------------------
best dev=0.000
holdout best=0.000 (baseline=0.000, Δ+0.000) → winner kept
proposer failed 2x — LLMProviderError: OpenAI API error 404 …
Persistence & the UI¶
When optimize(..., persist=True) (the default), the run is recorded to the
local local.db and surfaces in fastaiagent ui under AutoLLM — no
extra wiring. Two tables hold the record:
optimize_runs— one parent row per run: baseline/best dev scores, the holdout-guard scores,stopped_reason,reverted, theseed, the activelevers, the winningCandidateas JSON (for reproducibility), and inmetadatathe agent's original system prompt (baseline_system_prompt) and anyproposer_errors.optimize_iterations— one row per trajectory point:iteration,lever,dev_score,accepted/skipped,rationale, and aneval_run_id.
The eval_run_id is the key to the drill-down. Every candidate is scored by
a real aevaluate(persist=…) call, so it already lands in eval_runs /
eval_cases with traced runs. The iteration row just links to that existing
eval run — optimize stores no duplicate eval data. In the UI you can follow:
The view is read-only and refresh-based (REST, no live streaming). Open a run to see:
- Summary — baseline and best dev scores, the holdout score, and why the run stopped.
- Winner — the winning system prompt next to the prompt the run started from (each with a copy button), the few-shot examples and learned facts it selected, and any proposer failures. A reverted run, or one where nothing beat the baseline, says the agent keeps its original configuration.
- Trajectory —
baseline → accepted/skipped steps → holdout-guarded winnerwith per-iteration lever attribution; click any row through to the eval that produced its score.
Each run targets a single agent, so with several agents the list shows one row
per run tagged by agent_name; an agent filter appears once more than one
agent has runs (backed by GET /api/optimizes?agent=…).
Persistence is gated by the same persist flag that gates per-candidate evals,
so optimize(..., persist=False) writes nothing to optimize_runs /
optimize_iterations (and skips the per-candidate eval_runs writes too).
CLI¶
fastaiagent optimize \
--agent myapp.py:agent \
--dataset cases.jsonl \
--scorers exact_match \
--max-iterations 5 \
--levers instructions,fewshot \
--judge "is the answer correct and concise" \
--audit-judge "is the answer correct, complete and concise" \
--out winning_prompt.txt
--agentis apath/to/file.py:attrorpkg.module:attrthat resolves to anAgent.--leversis a comma-separated subset ofinstructions,fewshotandmemory(defaultinstructions).--judgeadds an LLM judge (a criteria string) as the selection scorer;--audit-judgeadds a distinct one used only on the holdout guard.--outwrites the winning system prompt before the summary is printed.
When not to use it¶
- Tiny datasets (< ~15 cases) can't form a meaningful 3-way split — run
harden()once instead.optimize()warns below 15 and errors below 3. VectorBlock-bearing agents can't be isolated per candidate (the block writes to an external store mid-run) —optimize()refuses unless you passallow_writable_memory=True. Other memory blocks are isolated automatically.- Tool/retrieval-bound agents — if quality is dominated by tool correctness rather than the prompt, fix the tools first.
- A
Supervisor,SwarmorChain—optimize()takes anAgentand raisesTypeErrorfor anything else. Optimize the agent behind each step.
Cost¶
The bill compounds: iterations × candidates × dev-size × judge-calls. Two hard
caps bound it:
max_eval_runscounts every evaluation pass: the baseline, each train re-score, each candidate, the few-shot teacher pass, and the holdout guard's passes.max_judge_callscounts one call per case for every model-backed scorer in a pass —LLMJudge/GEval,DecisionJudge, and the built-in RAG, agent, session and safety metrics — whether passed inscorersor asselection_judge/audit_judge. A customScorerthat calls a model itself isn't counted.
The loop holds back what the holdout guard needs, so the guard always runs and
neither total ever passes its cap. Caps too small for the baseline plus the guard
raise ValueError before anything runs. Select on a cheap deterministic scorer,
reserve the LLM judge for the holdout audit, and let patience / min_delta
stop early on noise.