LLM Judge¶
Use an LLM to evaluate output quality. The judge LLM scores agent responses based on criteria you define.
Basic Usage¶
from fastaiagent.eval import LLMJudge, evaluate
from fastaiagent import LLMClient
judge = LLMJudge(
criteria="correctness",
llm=LLMClient(provider="openai", model="gpt-4.1"),
)
# Use in evaluation
results = evaluate(
agent_fn=my_agent,
dataset=dataset,
scorers=[judge],
)
Custom Judge Prompt¶
Provide a custom prompt template to control exactly how the judge evaluates:
judge = LLMJudge(
criteria="helpfulness",
prompt_template=(
"Rate the following response for helpfulness.\n\n"
"User question: {input}\n"
"Expected answer: {expected}\n"
"Actual response: {output}\n\n"
'Respond with JSON: {{"score": <0.0-1.0>, "reasoning": "<explanation>"}}'
),
llm=LLMClient(provider="anthropic", model="claude-sonnet-4-6"),
)
The judge LLM must respond with JSON containing score and reasoning fields.
Placeholders¶
The canonical placeholders are {input}, {output} and {expected}. For compatibility with
templates authored on the control plane, these legacy aliases are also accepted and interpolate
identically:
| Placeholder | Aliases accepted |
|---|---|
{input} |
{{input}} |
{output} |
{{output}} |
{expected} |
{expected_output}, {{expected}}, {{expected_output}} |
Prefer the canonical single-brace spelling in new templates. Interpolation is a plain
substitution, not str.format(), so a literal JSON example such as {"score": ...} in your
template is left untouched — and content substituted into the template is never re-scanned,
so an input that happens to contain the text {output} stays literal.
If a custom prompt_template contains none of the known placeholders, the judge never sees the
content it is scoring, so a UserWarning is raised at construction time.
A template pulled with Scorer.from_platform("my-judge") is served verbatim by the platform and
follows the same rules — whichever spelling it was authored in.
Verdict parsing¶
Judge models sometimes wrap their verdict in prose or truncate the JSON. Rather than silently
scoring 0.0, the judge:
- parses the reply as JSON (tolerating ```json code fences);
- falls back to a regex scan for
"score"/"reasoning"if that fails; - retries once, handing the model its own reply back and demanding bare JSON;
- only then returns
score=0.0, passed=Falsewith areasonthat quotes the unparseable reply.
A transport/LLM error reports a different reason (Judge error: …) than an unreadable verdict
(Judge verdict unparseable after retry: …), so the two are distinguishable in results.
Scale and threshold¶
The raw verdict is normalized from scale into 0–1, and passed = score >= threshold
(default 0.5). With scale="1-5", a judge verdict of 4 becomes 0.75.
G-Eval (evaluation steps + rubric)¶
For richer, more reliable judging, pass evaluation_steps and/or a score-band rubric. This turns the judge into a G-Eval: it reasons step-by-step through your evaluation steps, scores against the rubric, and normalizes the result to 0–1. The plain criteria-only judge above is unchanged — G-Eval activates only when you provide steps or a rubric (or use the GEval class).
from fastaiagent.eval import GEval
judge = GEval(
name="correctness",
criteria="Is the answer factually correct and complete?",
evaluation_steps=[
"Identify the factual claim the answer makes.",
"Compare it against the expected answer.",
"Penalize fabricated, missing, or contradicted facts.",
],
rubric=[
(1, "Mostly incorrect"),
(3, "Partially correct"),
(5, "Fully correct"),
],
scale="1-5",
threshold=0.6, # on the normalized 0–1 score
)
result = judge.score(input="Capital of France?", output="Paris", expected="Paris")
print(result.score, result.passed, result.reason)
GEval is a thin, DeepEval-familiar wrapper over LLMJudge's G-Eval mode — these are equivalent:
from fastaiagent.eval import GEval, LLMJudge
GEval(name="x", criteria="...", evaluation_steps=[...], rubric=[...])
LLMJudge(criteria="...", evaluation_steps=[...], rubric=[...], scale="1-5", name="x")
Each instance can carry its own name, so several judges don't collide in the results (e.g. GEval(name="correctness") and GEval(name="tone")).
Auto-generated steps (Auto-CoT)¶
Give GEval a criteria but no evaluation_steps and it generates them from the criteria on first use (one extra LLM call, cached on the instance):
judge = GEval(name="helpfulness", criteria="Does the response directly answer the question?")
judge.score(input=question, output=answer)
print(judge.evaluation_steps) # derived steps, cached
The rubric is a list of (score_value, description) anchors on your scale; the judge interpolates between them and the final score is normalized to 0–1, with passed = score >= threshold.
See examples/81_g_eval.py for a runnable end-to-end script.
Scale Types¶
scale sets the range the judge scores on; the raw score is then normalized to 0–1. It applies on the G-Eval path (with evaluation_steps/rubric, or GEval); the legacy criteria-only judge always scores 0–1.
GEval(name="q", criteria="quality", scale="binary") # 0 or 1
GEval(name="q", criteria="quality", scale="0-1") # 0.0 to 1.0
GEval(name="q", criteria="quality", scale="1-5") # 1 to 5, normalized to 0–1
DecisionJudge (Decisions API)¶
New in 1.84.0. DecisionJudge asks OpenAI's Decisions API
instead of a chat model. It gets a probability back, so there's no verdict to
parse, no retry on an unreadable reply, and the score is calibrated rather than
self-reported.
- Predicate (default):
criteriais a statement that should be true of a good output, and the score is its probability. - Score: pass
levels, worst first. The score is the expected level, normalized to 0..1.
from fastaiagent.eval import DecisionJudge
judge = DecisionJudge("The actual output answers the input correctly.")
judge.score(input="What is 2+2?", output="4", expected="4") # score=1.00 passed=True
judge.score(input="What is 2+2?", output="5", expected="4") # score=0.00 passed=False
graded = DecisionJudge(
"How correct is the actual output?",
levels=["Wrong", "Partially correct", "Correct"],
threshold=0.75,
)
passed = score >= threshold(default 0.5). It drops intoevaluate(), the pytest gates andsimulate()anywhere anLLMJudgegoes.llmdefaults toLLMClient(model="gpt-6-luna").templatetakes the same{input}/{output}/{expected}placeholders, and their aliases, asLLMJudge. The rendered case is sent as evidence, never mixed into the question.- A refusal or a failed call scores
0.0with the reason spelled out, never a pass. - A judge that couldn't judge raises at construction. That covers empty criteria, a single level, a template with no placeholder, and a threshold outside 0..1.
- In
simulate(), the judge is re-aimed at each success criterion. Failure criteria are asked as "this undesirable condition occurred" predicates at 0.5.
Combining with Other Scorers¶
LLM judges work alongside built-in and custom scorers:
from fastaiagent.eval import evaluate
from fastaiagent.eval.builtins import ExactMatch, LengthBetween
results = evaluate(
agent_fn=my_agent.run,
dataset=dataset,
scorers=[
ExactMatch(),
LengthBetween(min_len=20, max_len=500),
LLMJudge(criteria="helpfulness"),
LLMJudge(criteria="correctness"),
],
)
Next Steps¶
- Evaluation — Core evaluation documentation
- Trajectory Scoring — Evaluate the path an agent took
- Session Scoring — Evaluate multi-turn conversations