Skip to content

LLM Judge

Use an LLM to evaluate output quality. The judge LLM scores agent responses based on criteria you define.

Basic Usage

from fastaiagent.eval import LLMJudge, evaluate
from fastaiagent import LLMClient

judge = LLMJudge(
    criteria="correctness",
    llm=LLMClient(provider="openai", model="gpt-4.1"),
)

# Use in evaluation
results = evaluate(
    agent_fn=my_agent,
    dataset=dataset,
    scorers=[judge],
)

Custom Judge Prompt

Provide a custom prompt template to control exactly how the judge evaluates:

judge = LLMJudge(
    criteria="helpfulness",
    prompt_template=(
        "Rate the following response for helpfulness.\n\n"
        "User question: {input}\n"
        "Expected answer: {expected}\n"
        "Actual response: {output}\n\n"
        'Respond with JSON: {{"score": <0.0-1.0>, "reasoning": "<explanation>"}}'
    ),
    llm=LLMClient(provider="anthropic", model="claude-sonnet-4-6"),
)

The judge LLM must respond with JSON containing score and reasoning fields.

Placeholders

The canonical placeholders are {input}, {output} and {expected}. For compatibility with templates authored on the control plane, these legacy aliases are also accepted and interpolate identically:

Placeholder Aliases accepted
{input} {{input}}
{output} {{output}}
{expected} {expected_output}, {{expected}}, {{expected_output}}

Prefer the canonical single-brace spelling in new templates. Interpolation is a plain substitution, not str.format(), so a literal JSON example such as {"score": ...} in your template is left untouched — and content substituted into the template is never re-scanned, so an input that happens to contain the text {output} stays literal.

If a custom prompt_template contains none of the known placeholders, the judge never sees the content it is scoring, so a UserWarning is raised at construction time.

A template pulled with Scorer.from_platform("my-judge") is served verbatim by the platform and follows the same rules — whichever spelling it was authored in.

Verdict parsing

Judge models sometimes wrap their verdict in prose or truncate the JSON. Rather than silently scoring 0.0, the judge:

  1. parses the reply as JSON (tolerating ```json code fences);
  2. falls back to a regex scan for "score" / "reasoning" if that fails;
  3. retries once, handing the model its own reply back and demanding bare JSON;
  4. only then returns score=0.0, passed=False with a reason that quotes the unparseable reply.

A transport/LLM error reports a different reason (Judge error: …) than an unreadable verdict (Judge verdict unparseable after retry: …), so the two are distinguishable in results.

Scale and threshold

The raw verdict is normalized from scale into 0–1, and passed = score >= threshold (default 0.5). With scale="1-5", a judge verdict of 4 becomes 0.75.

G-Eval (evaluation steps + rubric)

For richer, more reliable judging, pass evaluation_steps and/or a score-band rubric. This turns the judge into a G-Eval: it reasons step-by-step through your evaluation steps, scores against the rubric, and normalizes the result to 0–1. The plain criteria-only judge above is unchanged — G-Eval activates only when you provide steps or a rubric (or use the GEval class).

from fastaiagent.eval import GEval

judge = GEval(
    name="correctness",
    criteria="Is the answer factually correct and complete?",
    evaluation_steps=[
        "Identify the factual claim the answer makes.",
        "Compare it against the expected answer.",
        "Penalize fabricated, missing, or contradicted facts.",
    ],
    rubric=[
        (1, "Mostly incorrect"),
        (3, "Partially correct"),
        (5, "Fully correct"),
    ],
    scale="1-5",
    threshold=0.6,   # on the normalized 0–1 score
)

result = judge.score(input="Capital of France?", output="Paris", expected="Paris")
print(result.score, result.passed, result.reason)

GEval is a thin, DeepEval-familiar wrapper over LLMJudge's G-Eval mode — these are equivalent:

from fastaiagent.eval import GEval, LLMJudge

GEval(name="x", criteria="...", evaluation_steps=[...], rubric=[...])
LLMJudge(criteria="...", evaluation_steps=[...], rubric=[...], scale="1-5", name="x")

Each instance can carry its own name, so several judges don't collide in the results (e.g. GEval(name="correctness") and GEval(name="tone")).

Auto-generated steps (Auto-CoT)

Give GEval a criteria but no evaluation_steps and it generates them from the criteria on first use (one extra LLM call, cached on the instance):

judge = GEval(name="helpfulness", criteria="Does the response directly answer the question?")
judge.score(input=question, output=answer)
print(judge.evaluation_steps)   # derived steps, cached

The rubric is a list of (score_value, description) anchors on your scale; the judge interpolates between them and the final score is normalized to 0–1, with passed = score >= threshold.

See examples/81_g_eval.py for a runnable end-to-end script.

Scale Types

scale sets the range the judge scores on; the raw score is then normalized to 0–1. It applies on the G-Eval path (with evaluation_steps/rubric, or GEval); the legacy criteria-only judge always scores 0–1.

GEval(name="q", criteria="quality", scale="binary")   # 0 or 1
GEval(name="q", criteria="quality", scale="0-1")      # 0.0 to 1.0
GEval(name="q", criteria="quality", scale="1-5")      # 1 to 5, normalized to 0–1

DecisionJudge (Decisions API)

New in 1.84.0. DecisionJudge asks OpenAI's Decisions API instead of a chat model. It gets a probability back, so there's no verdict to parse, no retry on an unreadable reply, and the score is calibrated rather than self-reported.

  • Predicate (default): criteria is a statement that should be true of a good output, and the score is its probability.
  • Score: pass levels, worst first. The score is the expected level, normalized to 0..1.
from fastaiagent.eval import DecisionJudge

judge = DecisionJudge("The actual output answers the input correctly.")
judge.score(input="What is 2+2?", output="4", expected="4")   # score=1.00 passed=True
judge.score(input="What is 2+2?", output="5", expected="4")   # score=0.00 passed=False

graded = DecisionJudge(
    "How correct is the actual output?",
    levels=["Wrong", "Partially correct", "Correct"],
    threshold=0.75,
)
  • passed = score >= threshold (default 0.5). It drops into evaluate(), the pytest gates and simulate() anywhere an LLMJudge goes.
  • llm defaults to LLMClient(model="gpt-6-luna").
  • template takes the same {input} / {output} / {expected} placeholders, and their aliases, as LLMJudge. The rendered case is sent as evidence, never mixed into the question.
  • A refusal or a failed call scores 0.0 with the reason spelled out, never a pass.
  • A judge that couldn't judge raises at construction. That covers empty criteria, a single level, a template with no placeholder, and a threshold outside 0..1.
  • In simulate(), the judge is re-aimed at each success criterion. Failure criteria are asked as "this undesirable condition occurred" predicates at 0.5.

Combining with Other Scorers

LLM judges work alongside built-in and custom scorers:

from fastaiagent.eval import evaluate
from fastaiagent.eval.builtins import ExactMatch, LengthBetween

results = evaluate(
    agent_fn=my_agent.run,
    dataset=dataset,
    scorers=[
        ExactMatch(),
        LengthBetween(min_len=20, max_len=500),
        LLMJudge(criteria="helpfulness"),
        LLMJudge(criteria="correctness"),
    ],
)

Next Steps