Evals · 6 min

Build an AI agent eval set from real tickets

Build an AI agent eval set from real tickets. Stratify volume and failures, label correct answers, gate deploys. Book the $25,000 AI Transformation Sprint.

Feature graphic: Build an AI agent eval set from real tickets

To build an AI agent eval set from real tickets, export production history for one workflow, stratify by volume and failure mode, and have a domain owner label what "correct" means before any model change reaches traffic. Invented prompts are not a suite. In a published D2C support deployment, testing against 2,000 historical tickets with known resolutions produced 94.2% intent classification accuracy, 91.7% response accuracy, 99.1% policy compliance, and a 4.3/5.0 tone score before staged rollout. Which numbers should gate a deploy is a separate checklist. This post is the labeling workflow that produces those numbers — export, sample, define correct, and know when the set is good enough to block a ship.

Start from history, not from a brainstorm

Pull the last 90 days of tickets, forms, or call transcripts for one workflow. Keep the raw text the system will see: subject, body, structured fields, and the resolution that already happened. Do not rewrite tickets into "cleaner" demo language. The agent will meet the messy version in production.

A published support build began with classification of 28,400 tickets from 6 months of helpdesk history, then scored the agent against a 2,000-ticket slice with known resolutions. That sequence matters. Assessment maps the distribution. The eval set is the scored subset you refuse to ship without. If you only have a dozen happy-path scripts from a workshop, you have a demo script — not a ruler.

What to export on day one

  • Input as seen — raw ticket text, channel, and any fields the agent will receive.
  • Human resolution — the answer sent, the tool action taken, or the escalation reason.
  • Category / intent — whatever your team already tags, even if noisy.
  • Outcome flags — refund issued, return opened, policy exception, reopened ticket.

Strip customer PII before it leaves your environment. Keep enough structure that a labeler can still decide intent and the correct action. If legal blocks export, label inside the helpdesk with a locked view — do not invent substitute tickets to "move faster."

Stratify by volume and by the cases that hurt

A random sample of the top three intents will pass a demo and fail on the refund that creates a chargeback. Stratify on two axes: share of volume and cost of a miss.

BucketHow to sample itWhy it belongs
Top intents by volumeOversample the categories that burn hoursProtects the common path
Money / policy edgeReturns, discounts, cancellations, medical-advice refusalMisses here cost cash or compliance
Escalation-requiredCases a senior always takesTests whether the agent knows to stop
Adversarial / uglyTruncated text, rage, mixed languages, attachmentsCatches the tail demos skip
Recent policy changesTickets after a policy editStops the suite from scoring yesterday's rules

Aim for hundreds first — 200–500 labeled cases — then expand toward low thousands when volume supports it. The published support suite used 2,000 historical tickets before traffic moved. Under-sampling the awkward 10% is how you report 90% accuracy on the happy path and discover the rest from customers. For multi-step agents, score each irreversible step on its own labeled slice — the failure math is in when a multi-step AI agent fails in production.

Define "correct" before anyone scores a model

The suite is a specification. Write the rubric while the model is still offline. Domain owners decide; engineers wire the runner. If only the vendor labels correct, you bought a vibe check.

Fields every labeled case must carry

  • Intent / task label — the route the agent must choose.
  • Known-correct answer or action — the reply a senior would send, or the tool call that should fire.
  • Must-escalate flag — cases where autonomy is wrong even if the draft looks fine.
  • Policy tags — return window, discount eligibility, refusal rules.
  • Difficulty band — easy / edge / adversarial, so you never hide the tail inside one blended score.

Who labels, and how disagreements die

Pick one accountable owner per workflow — support lead, ops manager, clinical ops — not a rotating Slack channel. Have a second reviewer dual-label a 10–20% sample. Where they disagree, write the rule into the rubric and relabel. Unresolved disagreements become "must escalate" until the business picks a side. That disagreement rate is a product signal: if seniors cannot agree, the agent cannot either.

Do not let engineers invent the gold answer from memory of a demo. The eval suite is the strategy only when correct is a business decision with a name next to it.

Know when the set is good enough to gate a deploy

A set is ready when three conditions hold — not when someone is bored of labeling.

1. Coverage — top volume intents and the expensive failure modes each have enough labeled cases to move a percentage point. 2. Rubric stability — dual-label disagreements on the sample are resolved, and new tickets rarely invent a new "correct" mid-week. 3. Runner exists — the same cases score automatically on every prompt, model, and retrieval change, and a red run blocks promote.

Then write the floors into the SOW. The published support pattern — not a universal law — hit 94.2% intent, 91.7% response accuracy, 99.1% policy compliance, and 4.3/5.0 tone, with online autonomy only above 0.85 intent confidence and 0.80 retrieval relevance. Put your floors next to those numbers. Softening them after a bad demo week is how pilots never exit. After launch, promote hard live traces into the set weekly — especially escalations and low-confidence paths — so the suite tracks the traffic you actually get.

Failure modes that fake a suite

Workshop prompts. Marketing scenarios with perfect grammar. They pass every vendor demo and miss the truncated mobile ticket. Vendor-only labels. The builder marks its own homework. Scores rise; production does not. Happy-path only. Ninety percent of volume, zero chargeback cases. The blended score looks green. One blended metric. Intent is fine; the write-back is wrong. Without per-step labels you ship the product of the misses. Frozen set. Policy changed in July; the suite still scores June. Green CI, wrong answers.

If a vendor cannot run your labeled tickets, they cannot ship your agent. Refuse a demo-only milestone.

What to do this week

1. Export 90 days of one workflow's tickets with resolutions. Keep raw input; strip PII. 2. Stratify — top intents by volume, plus money/policy edges, mandatory escalations, and ugly cases. 3. Label 200–500 with a named domain owner: intent, correct action, must-escalate, policy tags, difficulty. 4. Dual-label 10–20% and resolve disagreements into the rubric before you score a model. 5. Price the diagnostic — book the $25,000, 2-week AI Transformation Sprint to turn that export into a suite and a scoped pilot that credits toward a build, or start from Gigabit Agents (from $8,000 flat) when the workflow and labels are already clear.

You cannot gate what you have not labeled. Build the eval set from the tickets your team already resolved, put a name on "correct," and treat every model swap as a deploy that needs a green run. The support proof shows the 2,000-ticket suite and the scores before traffic moved — and the metrics checklist is what those scores should look like in the contract.

Evals · FAQ

Questions this raises

How do you build an AI agent eval set from real tickets?

Export production tickets for one workflow with the raw input and the human resolution, stratify by volume and by costly failure modes, and have a domain owner label intent, the known-correct answer or action, must-escalate cases, and policy tags. Start with 200–500 labeled cases and expand toward low thousands when volume supports it. A published support deployment scored against 2,000 historical tickets with known resolutions before staged rollout. Invented workshop prompts are not an eval set.

Who should label correct answers on an AI agent eval set?

A named domain owner for that workflow — support lead, ops manager, or equivalent — not only the vendor and not a rotating Slack channel. Engineers wire the scoring runner. Dual-label a 10–20% sample; resolve disagreements into the rubric or mark the case must-escalate until the business decides. If seniors cannot agree on correct, the agent cannot either.

How do you know an AI agent eval set is ready to gate a deploy?

When top volume intents and expensive failure modes are covered, the labeling rubric is stable after dual-review, and an automated runner can score every prompt, model, and retrieval change with a red run that blocks promote. Then write pass floors into the SOW. In a published D2C support build, offline evals hit 94.2% intent accuracy, 91.7% response accuracy, 99.1% policy compliance, and 4.3/5.0 tone before traffic moved.

Why not invent prompts instead of labeling real tickets?

Invented prompts match the demo, not production. They miss truncated text, policy edges, mandatory escalations, and the ugly 10% that creates refunds and chargebacks. Teams that skip real tickets discover regressions from customers instead of from a red CI run. Pull history, label what already happened, and promote hard live traces into the set after launch.

Keep reading

Related insights

Evals

What to measure in an AI agent eval suite

AI agent eval metrics that should gate every deploy: golden-set size, accuracy thresholds, confidence gates.…

Evals

You don’t have an AI strategy until you have an eval suite

A model you can’t measure is a model you can’t trust in production. How we build evals before we build the a…

AI Agents

When a multi-step AI agent fails in production

A multi-step AI agent at 95% per step is about 66% end to end after eight steps. Split or gate it. Book the …

Stop reading, start shipping

Put a forward-deployed team on it.

If this is the kind of work you're trying to get into production, a 30-minute discovery call is the fastest path to a scoped plan.