Evals · 6 min

Calibrate an LLM-as-judge for AI agents

LLM-as-judge for AI agents is usable only after a human-labeled holdout calibration. Trust auto-scores then. Book the $25,000 AI Transformation Sprint.

Feature graphic: Calibrate an LLM-as-judge for AI agents

LLM-as-judge for AI agents is usable only after you calibrate it against a human-labeled holdout — not after the first rubric prompt looks plausible. Auto-scoring without that check is how teams greenlight a regression and call it quality. The ruler those labels produce is concrete: in a published D2C support deployment, a suite scored on 2,000 historical tickets with known resolutions hit 94.2% intent classification accuracy, 91.7% response accuracy, 99.1% policy compliance, and a 4.3/5.0 tone match before staged rollout. Which metrics belong in that suite is covered in what to measure in an AI agent eval suite. Why the suite comes before strategy slides is covered in you don't have an AI strategy until you have an eval suite. This post is the scoring layer in between: when an LLM may judge, when a human must, and when disagreement blocks the deploy.

Exact match vs LLM-as-judge

Split the suite by how "correct" is defined. If a domain owner can write a binary check, use the binary check. If the answer is free text with more than one acceptable phrasing, you need a judge — and a calibration loop before you trust it.

Score typeExamplePrefer
Exact / deterministicIntent label equals `return_status`; refund amount matches policy table; required escalation firedCode check, not a judge
Structured field matchTool call name and args match the labeled actionSchema compare + exact args
Free-text qualityReply matches known-correct meaning; tone matches brand baselineCalibrated LLM-as-judge
Irreversible judgmentMoney moved, clinical note filed, contract term waivedHuman label; judge drafts only

Policy compliance in the support proof is a 99.1% hit rate against documented rules. That is closer to exact match than to vibes: the rule either held or it did not. Response accuracy at 91.7% is the free-text band — different correct phrasings still count as correct. Put the first band in code. Put the second band behind a judge only after agreement on a holdout. Treating every metric as free text is how you pay for a second model to restate what a string compare already knew.

Calibrate on a human-labeled holdout

Calibration means a fixed set of cases where a human already scored the agent output, then you measure how often the judge agrees. Promote auto-scoring only when that agreement clears a floor you wrote down. Until then, the judge is a draft scorer, not a deploy gate.

Build the holdout like a ruler, not a demo

Pull real agent outputs — including failures — not invented chat. Use the same domain owners who labeled the golden set. Keep the holdout frozen while you tune the judge prompt; if you keep editing labels to make the judge look better, you measured nothing.

An instructional agreement bar

Assume this pattern unless your risk profile is stricter: do not promote LLM-as-judge as the primary score until judge–human agreement on the holdout clears about 85%, and dig every disagreement before you raise the bar. That 85% is an operator assumption for this guide, not a client measurement and not a substitute for the offline floors in the support proof (94.2% / 91.7% / 99.1%). Those floors are agent quality against human-known-correct labels. Judge agreement is whether your auto-scorer tracks those humans. Confusing the two numbers is the failure mode.

What to log on every disagreement

  • The case id and the human score.
  • The judge score and the one-line reason it gave.
  • Whether the rubric was ambiguous, the human was wrong, or the judge hallucinated a rule.
  • The fix: rubric edit, holdout re-label, or a hard exact check that should never have gone to a judge.

Re-run the frozen holdout after each rubric change. A green agent suite with a red judge–human agreement is not a pass — it is a broken measuring stick.

What the judge may score vs what humans must

Give the judge the work that is high volume and recoverable. Keep humans on the work that is rare, irreversible, or undefined.

Safe for a calibrated judge

Tone match against a written brand baseline (the support suite used 4.3/5.0 against the team's best agent). Semantic match of a reply to a known-correct answer when several phrasings are fine. Flagging likely policy language for a second pass — not the final compliance call when the rule is binary.

Keep humans on the holdout forever

Irreversible actions: refunds above a threshold, payment release, anything that writes money or legal status. Clinical-adjacent or regulated wording where a miss is silent until audit. Cases where two senior operators disagree — if humans cannot define correct, you are not past the first production gate.

Online confidence gates are a separate layer from judging. In the same support deployment, the agent went autonomous only when intent confidence exceeded 0.85 and retrieval relevance exceeded 0.80; below either threshold a human got the draft. Offline scoring and online gating both need numbers. Neither replaces the other. The production AI agent stack puts evals next to retrieval, observability, and guardrails for that reason.

Block the deploy on disagreement

Wire three checks into the same CI path that already runs the golden set:

1. Agent suite green — intent, response, policy, and any other gate metrics you put in the SOW. 2. Judge–human holdout green — agreement at or above the floor you chose; sample disagreements reviewed. 3. No silent rubric drift — judge prompt and rubric version pinned; a change is a deploy event.

If the agent suite passes but the holdout agreement drops, fail the build. The model under test is not the only thing that can regress — the scorer can too. Vendor model updates that retune judge behavior are a forced recalibration, the same way a production agent change is a forced golden-set run.

After launch, promote hard live traces into both the golden set and the judge holdout. That rhythm is what Managed AI Operations covers at $3,000–$20,000/month: regressions as numbers, not as customer tickets. In the published support build, 64% of tickets ran autonomous (3,149/month) only after the offline suite passed and confidence gates held — not after a judge prompt was pasted into a notebook.

What to do this week

1. Split your suite into exact checks vs free-text scores. Move every binary policy and schema match out of the judge. 2. Carve a frozen human-labeled holdout — start with dozens of scored outputs if volume is low; expand with the nasty tail. Do not invent chat. 3. Write the agreement floor on one line (use ~85% as a starting assumption, or raise it for irreversible workflows) and the agent pass floors next to it. 4. Fail CI when holdout agreement drops or the agent suite goes red. Pin judge prompt versions like any other deploy artifact. 5. Price the diagnostic: book the $25,000, 2-week AI Transformation Sprint to build the suite, the holdout, and a scoped pilot path that credits toward a build — or start from Gigabit Agents (from $8,000 flat) when the workflow and labels are already clear.

You cannot auto-score what you have not calibrated. Put human labels on a holdout, measure judge agreement, and keep exact checks in code. The full support proof shows what a human-known-correct suite looked like before traffic moved — 94.2% / 91.7% / 99.1% on 2,000 tickets, then 64% autonomous with gates at 0.85 and 0.80. Use that shape. Do not skip the calibration step and call the judge the truth.

Evals · FAQ

Questions this raises

When should you use an LLM-as-judge for AI agents?

Use a calibrated LLM-as-judge for free-text quality — semantic match to a known-correct answer, tone against a brand baseline — after a human-labeled holdout shows the judge agrees with domain owners. Prefer exact or schema checks for intent labels, tool-call args, and binary policy rules. Do not put irreversible money or regulated sign-off behind an uncalibrated judge.

What judge–human agreement rate should gate auto-scoring?

Write a floor before you promote auto-scoring. An instructional starting bar is about 85% agreement on a frozen human-labeled holdout, with every disagreement reviewed — raise it for irreversible workflows. That 85% is an operator assumption, not a published client metric. Keep it separate from agent accuracy floors such as the published support suite's 94.2% intent, 91.7% response, and 99.1% policy scores.

How is LLM-as-judge calibration different from AI agent eval metrics?

Eval metrics define what to measure and which numbers gate a deploy — intent, response accuracy, policy, escalation, tone. LLM-as-judge calibration defines whether an auto-scorer is allowed to produce those free-text scores at scale. You still need the golden set and pass thresholds; the judge is only the measuring stick for the free-text band after it tracks humans on a holdout.

Should we book the Sprint or Gigabit Agents to set this up?

Book the $25,000, 2-week AI Transformation Sprint when the workflow, golden set, and judge holdout are not yet defined — the Sprint maps the suite and a scoped pilot that credits toward a build. Buy Gigabit Agents from $8,000 flat when one production workflow already has labeled cases and clear gates. Managed AI Operations at $3,000–$20,000/month keeps regression runs and holdout refresh from slipping after launch.

Keep reading

Related insights

Evals

What to measure in an AI agent eval suite

AI agent eval metrics that should gate every deploy: golden-set size, accuracy thresholds, confidence gates.…

Evals

You don’t have an AI strategy until you have an eval suite

A model you can’t measure is a model you can’t trust in production. How we build evals before we build the a…

AI Agents

The production AI agent stack: what we actually deploy

Model, orchestration, retrieval, evals, observability, guardrails — the six layers every production agent ne…

Stop reading, start shipping

Put a forward-deployed team on it.

If this is the kind of work you're trying to get into production, a 30-minute discovery call is the fastest path to a scoped plan.