To build an AI agent eval set from real tickets, export production history for one workflow, stratify by volume and failure mode, and have a domain owner label what "correct" means before any model change reaches traffic. Invented prompts are not a suite. In a published D2C support deployment, testing against 2,000 historical tickets with known resolutions produced 94.2% intent classification accuracy, 91.7% response accuracy, 99.1% policy compliance, and a 4.3/5.0 tone score before staged rollout. Which numbers should gate a deploy is a separate checklist. This post is the labeling workflow that produces those numbers — export, sample, define correct, and know when the set is good enough to block a ship.
Start from history, not from a brainstorm
Pull the last 90 days of tickets, forms, or call transcripts for one workflow. Keep the raw text the system will see: subject, body, structured fields, and the resolution that already happened. Do not rewrite tickets into "cleaner" demo language. The agent will meet the messy version in production.
A published support build began with classification of 28,400 tickets from 6 months of helpdesk history, then scored the agent against a 2,000-ticket slice with known resolutions. That sequence matters. Assessment maps the distribution. The eval set is the scored subset you refuse to ship without. If you only have a dozen happy-path scripts from a workshop, you have a demo script — not a ruler.
What to export on day one
- Input as seen — raw ticket text, channel, and any fields the agent will receive.
- Human resolution — the answer sent, the tool action taken, or the escalation reason.
- Category / intent — whatever your team already tags, even if noisy.
- Outcome flags — refund issued, return opened, policy exception, reopened ticket.
Strip customer PII before it leaves your environment. Keep enough structure that a labeler can still decide intent and the correct action. If legal blocks export, label inside the helpdesk with a locked view — do not invent substitute tickets to "move faster."
Stratify by volume and by the cases that hurt
A random sample of the top three intents will pass a demo and fail on the refund that creates a chargeback. Stratify on two axes: share of volume and cost of a miss.
| Bucket | How to sample it | Why it belongs |
|---|---|---|
| Top intents by volume | Oversample the categories that burn hours | Protects the common path |
| Money / policy edge | Returns, discounts, cancellations, medical-advice refusal | Misses here cost cash or compliance |
| Escalation-required | Cases a senior always takes | Tests whether the agent knows to stop |
| Adversarial / ugly | Truncated text, rage, mixed languages, attachments | Catches the tail demos skip |
| Recent policy changes | Tickets after a policy edit | Stops the suite from scoring yesterday's rules |
Aim for hundreds first — 200–500 labeled cases — then expand toward low thousands when volume supports it. The published support suite used 2,000 historical tickets before traffic moved. Under-sampling the awkward 10% is how you report 90% accuracy on the happy path and discover the rest from customers. For multi-step agents, score each irreversible step on its own labeled slice — the failure math is in when a multi-step AI agent fails in production.
Define "correct" before anyone scores a model
The suite is a specification. Write the rubric while the model is still offline. Domain owners decide; engineers wire the runner. If only the vendor labels correct, you bought a vibe check.
Fields every labeled case must carry
- Intent / task label — the route the agent must choose.
- Known-correct answer or action — the reply a senior would send, or the tool call that should fire.
- Must-escalate flag — cases where autonomy is wrong even if the draft looks fine.
- Policy tags — return window, discount eligibility, refusal rules.
- Difficulty band — easy / edge / adversarial, so you never hide the tail inside one blended score.
Who labels, and how disagreements die
Pick one accountable owner per workflow — support lead, ops manager, clinical ops — not a rotating Slack channel. Have a second reviewer dual-label a 10–20% sample. Where they disagree, write the rule into the rubric and relabel. Unresolved disagreements become "must escalate" until the business picks a side. That disagreement rate is a product signal: if seniors cannot agree, the agent cannot either.
Do not let engineers invent the gold answer from memory of a demo. The eval suite is the strategy only when correct is a business decision with a name next to it.
Know when the set is good enough to gate a deploy
A set is ready when three conditions hold — not when someone is bored of labeling.
1. Coverage — top volume intents and the expensive failure modes each have enough labeled cases to move a percentage point. 2. Rubric stability — dual-label disagreements on the sample are resolved, and new tickets rarely invent a new "correct" mid-week. 3. Runner exists — the same cases score automatically on every prompt, model, and retrieval change, and a red run blocks promote.
Then write the floors into the SOW. The published support pattern — not a universal law — hit 94.2% intent, 91.7% response accuracy, 99.1% policy compliance, and 4.3/5.0 tone, with online autonomy only above 0.85 intent confidence and 0.80 retrieval relevance. Put your floors next to those numbers. Softening them after a bad demo week is how pilots never exit. After launch, promote hard live traces into the set weekly — especially escalations and low-confidence paths — so the suite tracks the traffic you actually get.
Failure modes that fake a suite
Workshop prompts. Marketing scenarios with perfect grammar. They pass every vendor demo and miss the truncated mobile ticket. Vendor-only labels. The builder marks its own homework. Scores rise; production does not. Happy-path only. Ninety percent of volume, zero chargeback cases. The blended score looks green. One blended metric. Intent is fine; the write-back is wrong. Without per-step labels you ship the product of the misses. Frozen set. Policy changed in July; the suite still scores June. Green CI, wrong answers.
If a vendor cannot run your labeled tickets, they cannot ship your agent. Refuse a demo-only milestone.
What to do this week
1. Export 90 days of one workflow's tickets with resolutions. Keep raw input; strip PII. 2. Stratify — top intents by volume, plus money/policy edges, mandatory escalations, and ugly cases. 3. Label 200–500 with a named domain owner: intent, correct action, must-escalate, policy tags, difficulty. 4. Dual-label 10–20% and resolve disagreements into the rubric before you score a model. 5. Price the diagnostic — book the $25,000, 2-week AI Transformation Sprint to turn that export into a suite and a scoped pilot that credits toward a build, or start from Gigabit Agents (from $8,000 flat) when the workflow and labels are already clear.
You cannot gate what you have not labeled. Build the eval set from the tickets your team already resolved, put a name on "correct," and treat every model swap as a deploy that needs a green run. The support proof shows the 2,000-ticket suite and the scores before traffic moved — and the metrics checklist is what those scores should look like in the contract.



