A multi-step AI agent fails in production when each step looks fine and the chain does not. If every step is right 95% of the time and the misses are independent, eight steps finish right about 66% of the time — 0.95 to the eighth is 0.663. Put that number in the review, not the model card. Shorten the chain, split it, or put a confidence gate on any step that writes money, a record, or a customer-visible message. What an agent costs is a separate budget. The layers we deploy are the production stack. Why a pilot stalls is the five gates. This post is only the step math.
Per-step accuracy is not the success rate
Quote the step, not the demo. A classifier at 95% on a labeled set is a one-step score. A workflow that classifies, retrieves, extracts a field, calls a tool, checks a policy, and writes back is six chances to be wrong. When those misses are independent, end-to-end accuracy is the product of the step rates. Four steps at 95% are already about 81%. Twelve are about 54%.
Independence is the assumption under the table, and it is not free. It holds when each step can miss for its own reason: a stale field, a tool timeout, a threshold that does not fire. It does not hold when one upstream miss — usually a bad retrieval — poisons every later step. In that case the product understates how often the whole run dies together, and a gate on the weak step beats another prompt on the last one. Say which case you are in before you multiply. These figures are rounded arithmetic. They are not a client measurement.
| Steps | 99% each | 95% each | 90% each |
|---|---|---|---|
| 1 | 99% | 95% | 90% |
| 4 | 96% | 81% | 66% |
| 8 | 92% | 66% | 43% |
| 12 | 89% | 54% | 28% |
95% of enterprise generative-AI pilots deliver no measurable P&L impact (MIT Project NANDA, 2025). Chain length is one mechanical reason a three-turn demo falls over on a twelve-call workflow. The pilot never had a product to fail. It had a slide that said 95%.
Count every step that can be wrong
A step is any place the run can pick the wrong thing. If you cannot list them, you do not have a success rate.
- Classify the request.
- Retrieve the policy, the order, or the record.
- Extract a field from messy text.
- Call a tool.
- Decide whether to proceed.
- Write back — a reply, a refund, a label, an address, a chart note.
The write-back is the step teams forget to score. They grade the sentence and ignore the tool that changes the system of record. That tool is the one with the blast radius.
The published counterexample is short on purpose. In a D2C support deployment, offline scoring on 2,000 historical tickets hit 94.2% intent accuracy, 91.7% response accuracy, and 99.1% policy compliance. If intent and response both had to be right, and if those two rates were independent, the joint figure would be about 86% (0.942 × 0.917). The case study does not publish that joint rate. It publishes the gates that kept the agent off the weak tail: an autonomous reply only when intent confidence exceeded 0.85 and retrieval relevance exceeded 0.80. Under that rule the agent resolved 64% of tickets with no human — 3,149 a month — not 94%. First response on those AI replies fell from 4.2 hours to 6 minutes. Year 1 savings were $217,200 against $92,000 of Year 1 investment. The 64% is the share that cleared the gates. It is not an eight-step product.
Copy the shape. A long chain without a gate ships the product of its misses. A short chain with a gate ships the slice it can defend. Which scores belong in the suite is what to measure before a deploy. Use that list on the steps you just counted. Do not invent a second metric taxonomy here.
Three moves that raise the end-to-end number
You have three levers. More model calls is not one of them.
Shorten the chain
Most of the path does not need a model. Eligibility checks, field validation, return-window tests, and label generation are deterministic. Put the model on the step that needs language — free-text classification, a messy vendor email — and keep the rest a state machine. Every model step you delete multiplies what remains by something closer to 1. The layer list for that state machine lives in the production stack.
Gate the irreversible step
Irreversible means a customer sees it, money moves, or a system of record changes. Those steps get a floor and a human. In the support build the floors were 0.85 on intent and 0.80 on retrieval. Below either one, a person received the draft and the context. Reversible reads — look up an order, draft a reply, classify a reason — can run hotter. The question per action is the one in what an agent can do without a person in the loop: how reversible is the action, and how fast would you notice a miss. A quiet miss that runs for a quarter costs more than a loud failure on day one, because nobody stops it.
Split one loop into two scored workflows
A twelve-step loop has one blended score and no owner. The average hides the last step. Two workflows — classify and retrieve, then act — each with a golden set, fail in a place you can see. The handoff is a typed object: intent, record id, confidence, the fields the next step is allowed to touch. If the second workflow cannot name its inputs, you still have one loop with extra branding.
Gigabit Agents are from $8,000 flat per agent for one production workflow. A twelve-step "one agent" is often two workflows and should be priced that way. Paying once for a chain that finishes right about 54% of the time, at 95% per step, is the expensive version.
Failure modes that survive the demo
Five patterns show up after the room claps.
The model card becomes the workflow rate. A slide says 95%. The workflow has eight tool calls. You bought about 66% and reported 95%.
One loop, no checkpoint. The agent may call tools until it decides it is finished. There is no step list, so there is no product to compute and no place to insert a gate.
A blended score. Say intent is 98% and the write-back is 80%. The average looks fine. The customer sees the write-back.
Autonomy set to the classification rate. 94.2% intent accuracy is not permission to auto-send 94% of traffic. The support proof auto-sent the slice above the gates and left the rest with people. Copy the gate, not the headline accuracy.
No labels, so no product. If you cannot put a measured rate on each step, stop multiplying. A vendor benchmark is not your step rate. Label a sample from real tickets first. That sample is also the start of the eval suite that separates a pilot from a system you can keep.
What to do this week
1. Write the chain for one workflow as a numbered list. Include retrieval, each tool call, the proceed-or-stop decision, and the write-back. 2. Put a rate on each step from a labeled sample of real cases. If you have no labels, that is the work — not another demo. 3. Multiply only if the misses are independent. If one upstream miss causes the rest, fix that step before you add length. The table above is the independent case, rounded. 4. Mark irreversible steps and set a floor. If you have no measured floor, do not auto-execute them. 0.85 intent and 0.80 retrieval are a published support example, not a default for payments or clinical notes. 5. Price the cut. Book the $25,000, two-week AI Transformation Sprint to map the chain, the gates, and a scoped build that credits toward an agent. If the workflow is already one clear step with a golden set, buy Gigabit Agents from $8,000 flat and skip the diagnostic.
A 95% step is not a 95% workflow. Count the steps, delete the ones a rule can do, and gate the ones that write. The published support proof is the reference: a short chain, two numeric gates, 64% of tickets handled with no human, and the rest handed over with a draft. The $25,000 Sprint is the two weeks to run that test on your own queue before you fund a loop that cannot clear 66%.



