Vendor Evaluation · 5 min

Nine questions to ask any AI agency before you sign

The questions that separate a firm that ships from one that demos — what each is really testing, and what a bad answer sounds like. Ask ours the same ones.

G

Most AI vendor evaluations are decided by a demo, which is close to the worst available signal. A demo proves a model can do a task once under conditions the vendor chose. You're buying the thing that does it ten thousand times under conditions they didn't. These are the nine questions that actually separate the two, what each one is testing underneath, and what the evasive version sounds like.

They matter because the base rate is poor: 95% of enterprise generative-AI pilots deliver no measurable P&L impact (MIT Project NANDA, 2025). The same study found the pattern that beats it — buying from or partnering with specialized vendors succeeded about 67% of the time, roughly three times the rate of internal builds. Partnering works. Partnering *badly* is most of the 95%.

1. What have you put into production, and what does it do now?

Not what they've built — what's *running*, today, with users depending on it. Ask for the workflow, the volume, and how long it's been live.

*Bad answer:* a portfolio of pilots, prototypes, and proofs of concept described in the past tense. Anyone can demo. Shipping and keeping it alive is the whole skill, and a firm that's done it will have specifics. Ours are on the proof page with the investment and the return on each.

2. Who exactly will do the work, and will I meet them before I sign?

The gap between the people in the pitch and the people on the keyboard is where most engagements go wrong. Ask for names, seniority, and whether they're on other accounts.

*Bad answer:* "our team," a org chart, or a partner who'll "stay close to the project." If the people who scope it aren't the people who build it, every piece of context you transfer gets transferred again — which is the entire argument for a forward-deployed model.

3. What's the total price, and who absorbs an overrun?

One number, in writing, with the answer to what happens when it goes long. This is the fastest question on the list for sorting engagement models.

*Bad answer:* a rate card, a range, or "it depends on scope." It always depends on scope; the question is who carries that risk. We wrote up the mechanism in Bait-and-Bill, and our own prices are published so you can hold us to the same standard.

4. How will we know it's working?

You're looking for a metric and a measurement method, ideally agreed before the build. Accuracy on what test set? Resolution rate against what baseline? Measured by whom, how often?

*Bad answer:* qualitative outcomes, adoption numbers, or "we'll define success together in discovery." A firm that builds AI for a living has opinions about how to measure it — that's what an eval suite is, and not having a view on yours is a tell.

5. What happens when the model gets it wrong?

Every probabilistic system fails. You're testing whether they've thought about the failure mode before you asked: guardrails, escalation paths, human checkpoints, and the blast radius on a bad day.

*Bad answer:* "the model is very accurate," or anything that treats reliability as a property of the model instead of the system around it. The real answer describes constraints in code, not confidence in a vendor's benchmarks.

6. What happens when the underlying model changes?

Model vendors deprecate, retune, and ship behavior changes on their own schedule. Ask what breaks, how they'd find out, and how fast they could roll back.

*Bad answer:* a blank look, or a claim of model-agnosticism with nothing behind it. The credible version involves pinned versions, versioned prompts, a regression suite, and a return path to a known-good state — the layers we describe in the production agent stack.

7. Who owns the code, the prompts, and the evals?

Ask explicitly about all three, plus the infrastructure and the data. Prompts and eval sets are the accumulated knowledge of the engagement, and they're the part most likely to be quietly retained.

*Bad answer:* ownership of "deliverables," or code ownership with the orchestration living on the vendor's platform. If leaving means rebuilding, the price of the engagement includes never leaving.

8. What does it cost to run after launch, and who runs it?

Live AI systems drift, and something has to watch them. Ask for the monthly number and who is accountable for it.

*Bad answer:* treating operations as optional or as a support SKU. Inference itself is cheap — cents per transaction, and falling fast — but monitoring, eval regressions, and drift management are real work. Ours is $3,000–$20,000 a month depending on how many systems are under management, and we'd rather quote it up front than discover it later.

9. When would you tell us not to build this?

The most revealing question on the list, and the one almost nobody asks. You're testing whether the firm has a view on fit at all, or whether every problem happens to be solvable by what they sell.

*Bad answer:* enthusiasm. A firm worth hiring will name the conditions under which you shouldn't proceed — the data isn't there, the workflow is too low-volume, the right answer isn't knowable, nobody owns the metric. If they've never talked a client out of something, you're not their client, you're their pipeline.

How to use the answers

Ask all nine of every vendor, including us, and write the answers down side by side. The pattern matters more than any single response: firms that ship give specific, slightly boring answers with numbers in them, and firms that demo give confident answers about capability. If you want the structural version of this comparison, we lay out the tradeoffs against traditional consultancies and automation agencies — including where each one is genuinely the better choice.

Vendor Evaluation · FAQ

Questions this raises

What questions should I ask an AI agency before hiring them?

Nine: what have you put into production and what does it do now; who exactly will do the work; what is the total price and who absorbs an overrun; how will we know it is working; what happens when the model gets it wrong; what happens when the underlying model changes; who owns the code, prompts, and evals; what does it cost to run after launch; and when would you tell us not to build this.

How do I tell a real AI engineering firm from a demo shop?

Ask what is running in production today, with volume and duration — not what has been built. Demo shops answer in the past tense with pilots and prototypes; firms that ship give specific, slightly boring answers with numbers in them and can describe their failure modes without being prompted.

What is the most revealing question to ask an AI vendor?

When would you tell us not to build this? It tests whether the firm has a view on fit at all. A vendor worth hiring will name the conditions under which you shouldn't proceed — thin data, low volume, no knowable right answer, no owner for the metric. Pure enthusiasm is the warning sign.

Should I ask who owns the prompts and eval suite?

Yes, explicitly, alongside the code, infrastructure, and data. Prompts and eval sets are the accumulated knowledge of the engagement and the part most often quietly retained. If leaving the vendor means rebuilding from scratch, the real price of the engagement includes never leaving.

Keep reading

Related insights

Vendor Evaluation

Bait-and-Bill: the AI pricing trap to watch for

The headline number isn't the price — it's the entry fee for an open-ended meter. How the trap is built, the…

Vendor Evaluation

Vendor Roulette: why switching AI vendors keeps you at square one

Third vendor, same starting line. Why each switch resets you to zero, what the previous engagement was suppo…

Vendor Evaluation

How to hand over an AI system — to a new vendor or in-house

The inventory to collect, the acceptance test that proves the handover worked, and the four things that alwa…

Stop reading, start shipping

Put a forward-deployed team on it.

If this is the kind of work you're trying to get into production, a 30-minute discovery call is the fastest path to a scoped plan.