Strategy

21 Questions to Ask an AI Agency Before You Sign

Jul 13, 20267 min read

Sequence beats the list. Most buyers open with cost, which teaches the vendor what to optimise the pitch for. Twenty-one questions in the order that works, grouped by scope, data, integration, ownership and failure, with what a good and a bad answer sound like.

21 Questions to Ask an AI Agency Before You Sign

Ask them in this order, and the price changes.

Most buyers open with cost and timeline. That single choice teaches the vendor what the pitch has to optimise for, and you spend the rest of the meeting hearing a number defended rather than a problem examined. The questions to ask an AI agency are more useful in a specific sequence: scope and discovery first, then data, integration, ownership and failure, with commercial last. Ask it that way and the price you eventually hear is priced against your actual constraints. Twenty-one questions follow, grouped, with the answers worth listening for.

Questions 1 to 4: what are we actually building

  • Which specific workflow will this system touch, and who performs it today?
  • What will you do in discovery, what does it cost, and do I own the output if I stop there?
  • What would make you tell me not to build this?
  • Which parts of my request would you cut, and why?

Question 3 is the most diagnostic sentence in the whole list. A good answer sounds like a firm describing a real case where they walked away or recommended an off-the-shelf tool, with enough detail that you can tell it happened. A bad answer sounds like enthusiasm: every idea is viable, every workflow is a candidate, and the only variable is budget. A vendor who has never talked a client out of anything is either very lucky or selling capacity.

Questions 5 to 12: where the project actually stalls

  • Where will my data be processed, in which region, and by which sub-processors?
  • Is my content used to train anybody’s model, and will you put that in the contract?
  • What retention applies to prompts and outputs, and can we shorten it?
  • Who on your side gets access to my production data, and how is that logged?
  • Which of my systems must this integrate with, and have you connected to them before?
  • What do you need from us, by when, and who blocks the project if it is late?
  • What happens when my source system changes its API?
  • How will this behave on our messiest real records, not the sample?

Question 10 predicts your timeline better than any Gantt chart. Projects stall on client-side access approvals and unavailable subject-matter experts far more often than on engineering. A good answer names the artefacts, the dates and the roles on your side, and treats them as part of the plan. A bad answer is “we’ll handle everything”, which sounds like service and means the dependency will surface in week six as your delay. If the data questions in group two are unfamiliar territory, where your data actually goes covers what to expect the answers to look like.

Question 9 is worth pressing twice, because “we integrate with everything” is technically true of anything with an API and tells you nothing about effort. A good answer names the system, the version, and something specific about it that only somebody who has connected to it would know: the rate limits, the awkward authentication, the field that is documented but empty in practice. A bad answer stays at the level of the logo. The follow-up that settles it is asking what surprised them last time they connected to it, since a supplier who has genuinely done the work always has a complaint ready.

Questions 13 to 18: what you hold when it ends

  • Who owns the code, the prompts, the fine-tunes and the evaluation sets?
  • What exactly is handed over, and in what state is the documentation?
  • Could my own developer run this without you in six months?
  • What does it cost to run per month once you have gone?
  • Which named people will do the work, and what else are they on?
  • What is your notice period, and what happens to the system on the day we part?

Question 17 exposes the pattern that costs buyers the most and has no polite name. Senior people sell the work; less experienced people deliver it. This is a structural risk of the category rather than a criticism of any firm, and it grows with the size of the supplier. A good answer gives you names, seniority and current allocation, and offers to put the key person in the contract. A bad answer describes a team in the abstract, or the person in front of you says “I’ll be across it” without committing to hours. The WA Center platform is a useful test case for question 14: it went live with role-based access for several distinct user types, and the WA Center case study sets out what was actually handed over, which is the level of specificity to ask any supplier for.

Question 13 needs its own attention because the wrong answer is expensive and easy to miss. Code ownership is the part buyers check; prompts, fine-tunes and evaluation sets are the parts they forget, and those are where a lot of the working knowledge sits. A good answer assigns all four to you on payment, in one clause, and the supplier can point to it without opening a laptop. A bad answer separates them: you own the code, they retain a licence to the prompts, or the evaluation set is described as their methodology. That structure is not always predatory, since some firms genuinely reuse internal tooling, but it means leaving is harder than you think and you should price that in rather than discover it later.

Questions 19 to 21: what happens when it does not work

  • What is the measurable definition of success, and what is the baseline today?
  • If we hit month three and the number has not moved, what happens contractually?
  • Which engagement went badly for you, and what did you change afterwards?

Question 19 catches the most common structural defect in AI proposals. If nobody records the baseline before the build, nobody can demonstrate the system worked, and the argument at month six becomes a matter of impressions. A good answer proposes the metric, admits which parts are hard to attribute, and offers to measure the baseline during discovery. A bad answer offers efficiency and productivity without a number attached to either.

Question 20 is the one most suppliers have never been asked, and their discomfort answering it is informative on its own. A good answer is a specific remedy: a rework period at their cost, a staged payment still outstanding at month three, or an honest statement that they carry no financial risk and here is why the price reflects that. A bad answer reframes the question as a partnership problem, or explains that outcomes depend on factors outside their control. Both of those may be true. Neither tells you what happens to your money.

Five red flags override everything above. A fixed quote arriving before any discovery. No named engineers. A pitch that leads with a demo rather than your workflow. No measurement plan. Intellectual property retained by the vendor by default. Any one of these is worth a direct challenge; two together, in our experience across the engagements we have run, usually means the proposal was written before anybody understood the problem.

Twenty-one questions is too many for a small job

Do not run this at a supplier quoting four thousand euros for a single automation. The procurement effort would exceed the contract value, and a good small vendor will reasonably decline to do that much paperwork. Take questions 3, 13 and 19 (what would you cut, who owns it, how will we know it worked) and sign. The full list earns its time somewhere north of roughly twenty thousand euros, or wherever a failure would be genuinely disruptive rather than merely annoying.

The list is also the wrong tool if you have not settled what you want. Twenty-one sharp questions aimed at a vague brief produce twenty-one confident answers to a problem nobody has defined, and you will have procured something precise and unhelpful. Fix the brief first: our analysis of why vendor relationships fail covers the contract-shaped version of that failure, and the audit process exists partly to produce a brief that is worth taking to market.

  • Sequence beats the list. Scope, data, integration, ownership and failure first; commercial last. Opening with price teaches the vendor what to optimise the pitch for.
  • The single most diagnostic question is what would make them tell you not to build. Firms that have said it before can describe when.
  • Ask what they need from you and by when. Client-side access approvals and unavailable experts stall more projects than engineering does.
  • Get the named people and their allocation in writing. Senior sells, junior delivers is a structural category risk, and it scales with supplier size.
  • Five red flags override the rest: fixed quote before discovery, no named engineers, demo-led pitch, no measurement plan, vendor-retained IP.

Every question here is yours to use with any supplier, including the ones you are comparing us against, and the list works better when the same twenty-one go to all of them. If you would rather arrive at the conversation with the brief already written, that is what an audit produces: a costed plan, a measurement baseline, and a specification any vendor can quote against. Which of the twenty-one would your current shortlist struggle with most?

Book your AI audit