LLM-as-judge for AI agents is usable only after you calibrate it against a human-labeled holdout — not after the first rubric prompt looks plausible. Auto-scoring without that check is how teams greenlight a regression and call it quality. The ruler those labels produce is concrete: in a published D2C support deployment, a suite scored on 2,000 historical tickets with known resolutions hit 94.2% intent classification accuracy, 91.7% response accuracy, 99.1% policy compliance, and a 4.3/5.0 tone match before staged rollout. Which metrics belong in that suite is covered in what to measure in an AI agent eval suite. Why the suite comes before strategy slides is covered in you don't have an AI strategy until you have an eval suite. This post is the scoring layer in between: when an LLM may judge, when a human must, and when disagreement blocks the deploy.
Exact match vs LLM-as-judge
Split the suite by how "correct" is defined. If a domain owner can write a binary check, use the binary check. If the answer is free text with more than one acceptable phrasing, you need a judge — and a calibration loop before you trust it.
| Score type | Example | Prefer |
|---|---|---|
| Exact / deterministic | Intent label equals `return_status`; refund amount matches policy table; required escalation fired | Code check, not a judge |
| Structured field match | Tool call name and args match the labeled action | Schema compare + exact args |
| Free-text quality | Reply matches known-correct meaning; tone matches brand baseline | Calibrated LLM-as-judge |
| Irreversible judgment | Money moved, clinical note filed, contract term waived | Human label; judge drafts only |
Policy compliance in the support proof is a 99.1% hit rate against documented rules. That is closer to exact match than to vibes: the rule either held or it did not. Response accuracy at 91.7% is the free-text band — different correct phrasings still count as correct. Put the first band in code. Put the second band behind a judge only after agreement on a holdout. Treating every metric as free text is how you pay for a second model to restate what a string compare already knew.
Calibrate on a human-labeled holdout
Calibration means a fixed set of cases where a human already scored the agent output, then you measure how often the judge agrees. Promote auto-scoring only when that agreement clears a floor you wrote down. Until then, the judge is a draft scorer, not a deploy gate.
Build the holdout like a ruler, not a demo
Pull real agent outputs — including failures — not invented chat. Use the same domain owners who labeled the golden set. Keep the holdout frozen while you tune the judge prompt; if you keep editing labels to make the judge look better, you measured nothing.
An instructional agreement bar
Assume this pattern unless your risk profile is stricter: do not promote LLM-as-judge as the primary score until judge–human agreement on the holdout clears about 85%, and dig every disagreement before you raise the bar. That 85% is an operator assumption for this guide, not a client measurement and not a substitute for the offline floors in the support proof (94.2% / 91.7% / 99.1%). Those floors are agent quality against human-known-correct labels. Judge agreement is whether your auto-scorer tracks those humans. Confusing the two numbers is the failure mode.
What to log on every disagreement
- The case id and the human score.
- The judge score and the one-line reason it gave.
- Whether the rubric was ambiguous, the human was wrong, or the judge hallucinated a rule.
- The fix: rubric edit, holdout re-label, or a hard exact check that should never have gone to a judge.
Re-run the frozen holdout after each rubric change. A green agent suite with a red judge–human agreement is not a pass — it is a broken measuring stick.
What the judge may score vs what humans must
Give the judge the work that is high volume and recoverable. Keep humans on the work that is rare, irreversible, or undefined.
Safe for a calibrated judge
Tone match against a written brand baseline (the support suite used 4.3/5.0 against the team's best agent). Semantic match of a reply to a known-correct answer when several phrasings are fine. Flagging likely policy language for a second pass — not the final compliance call when the rule is binary.
Keep humans on the holdout forever
Irreversible actions: refunds above a threshold, payment release, anything that writes money or legal status. Clinical-adjacent or regulated wording where a miss is silent until audit. Cases where two senior operators disagree — if humans cannot define correct, you are not past the first production gate.
Online confidence gates are a separate layer from judging. In the same support deployment, the agent went autonomous only when intent confidence exceeded 0.85 and retrieval relevance exceeded 0.80; below either threshold a human got the draft. Offline scoring and online gating both need numbers. Neither replaces the other. The production AI agent stack puts evals next to retrieval, observability, and guardrails for that reason.
Block the deploy on disagreement
Wire three checks into the same CI path that already runs the golden set:
1. Agent suite green — intent, response, policy, and any other gate metrics you put in the SOW. 2. Judge–human holdout green — agreement at or above the floor you chose; sample disagreements reviewed. 3. No silent rubric drift — judge prompt and rubric version pinned; a change is a deploy event.
If the agent suite passes but the holdout agreement drops, fail the build. The model under test is not the only thing that can regress — the scorer can too. Vendor model updates that retune judge behavior are a forced recalibration, the same way a production agent change is a forced golden-set run.
After launch, promote hard live traces into both the golden set and the judge holdout. That rhythm is what Managed AI Operations covers at $3,000–$20,000/month: regressions as numbers, not as customer tickets. In the published support build, 64% of tickets ran autonomous (3,149/month) only after the offline suite passed and confidence gates held — not after a judge prompt was pasted into a notebook.
What to do this week
1. Split your suite into exact checks vs free-text scores. Move every binary policy and schema match out of the judge. 2. Carve a frozen human-labeled holdout — start with dozens of scored outputs if volume is low; expand with the nasty tail. Do not invent chat. 3. Write the agreement floor on one line (use ~85% as a starting assumption, or raise it for irreversible workflows) and the agent pass floors next to it. 4. Fail CI when holdout agreement drops or the agent suite goes red. Pin judge prompt versions like any other deploy artifact. 5. Price the diagnostic: book the $25,000, 2-week AI Transformation Sprint to build the suite, the holdout, and a scoped pilot path that credits toward a build — or start from Gigabit Agents (from $8,000 flat) when the workflow and labels are already clear.
You cannot auto-score what you have not calibrated. Put human labels on a holdout, measure judge agreement, and keep exact checks in code. The full support proof shows what a human-known-correct suite looked like before traffic moved — 94.2% / 91.7% / 99.1% on 2,000 tickets, then 64% autonomous with gates at 0.85 and 0.80. Use that shape. Do not skip the calibration step and call the judge the truth.



