Is Your Data Ready for AI? An Honest Assessment
Your data is probably good enough for the first project, and almost certainly not for the one you have in mind. Readiness is a property of a use case, not of your company. Five questions, three tiers, and the cases where the honest fix is a form change.

Your data is probably good enough for the first project.
It is almost certainly not good enough for the project you have in mind, and those two statements are not in conflict. Data readiness for AI is a property of a specific use case, not a property of your company. A support assistant that answers from twelve policy documents needs almost nothing: the documents and permission to read them. A demand forecast across four warehouses needs three years of consistent history, and if your warehouse codes changed in 2024 you do not have it. Same company, same data, two completely different answers. Ask the question the second way and it becomes answerable in an afternoon.
Why a data readiness assessment stalls projects
The standard sequence runs: assess the data, clean the data, then build. It sounds responsible and it kills more projects than bad models do. A company-wide data programme takes nine to eighteen months, costs more than the AI project it was meant to enable, and produces a warehouse designed against requirements nobody has tested yet. Two years later the original sponsor has moved on and the AI initiative is a line item nobody defends. We have walked into that exact situation more than once, and the data was never the reason the project died. Sequencing was.
The inversion is straightforward. Pick the use case that fits the data you already have, ship it, and let the operational value fund the cleanup that the second use case needs. This is not a shortcut, it is a different order of operations, and it has one large advantage: after the first system runs for a quarter you know which data problems actually cost you money, rather than which ones look untidy in a schema diagram. Most cleanup backlogs are full of the second kind. The first system is the instrument that tells you which is which.
The data readiness check to run before you scope anything
Take one candidate use case. Not your whole company. Answer these about the specific data that use case needs.
- Where does it live, and how many places is that? One system is easy. Three is normal. When the answer is “it depends who you ask”, you have found the real project. Count the systems before anything else, because integration surface drives cost harder than model choice does.
- Who owns it? Not which department stores it. Which named person can approve access this week. Projects stall for a month here more often than they stall on anything technical, and the fix is a conversation, not a budget.
- Is it structured, or is it prose? Both are workable and they are workable for different things. Tables support forecasting and scoring. Documents, emails and transcripts support retrieval, summarising and classification. Companies routinely believe unstructured data is unusable, which stopped being true around 2023.
- Can a machine reach it? An API, a database connection, or an export somebody can schedule. If the honest answer is that a person copies it into a spreadsheet each Monday, that is not a blocker, it is the first thing worth automating and it pays for itself before the AI does.
- Is it accurate enough to act on? Not perfect. Enough. Pull fifty records at random and check them by hand. Anyone can do this in an hour, almost nobody does, and the result reorders the project plan more often than any workshop.
Score the answers and you land in one of three tiers. Ready means one or two systems, a named owner, reachable by machine, and a sample that holds up: build now. Needs work means the data exists and is reachable but the sample failed, or the owner is unclear: usually four to eight weeks of targeted fixing, scoped to this use case only. Not viable yet means the data does not exist, sits with a vendor who will not export it, or is so inconsistent that the sample is meaningless. That is the honest answer and it should send you to a different use case, not to a data programme.
Matching the use case to the data you actually have
Tolerant of messy data: retrieval and question answering over documents, classification and routing of inbound work, drafting that a human edits, extraction where a person confirms the result. These work because a human sits at the end of the process and the cost of an error is a correction rather than a bad decision. The Biscoito.ai assistant we built for a veterinary clinic runs on exactly this shape: existing clinic material, no data warehouse, a human in the loop for anything clinical. The Biscoito.ai case study shows what the input actually looked like, and it was not tidy.
Intolerant of messy data: forecasting, pricing, anything that scores or ranks people, and any system whose output goes straight into an operational decision without review. These need consistent history and stable definitions. If your product codes were renumbered, your regions were redrawn, or two systems disagree about what a customer is, a forecast will produce a confident number that is wrong, which is worse than no number.
On cost: targeted cleanup for a single use case is usually a few weeks of one person’s time, and in published European market terms sits at the lower end of the €3,000 to €50,000 band that custom AI work occupies, rather than being a separate programme. Those are market ranges, not our price list. The number that matters more is elapsed time: access approvals, not engineering, are what stretch a four-week fix into three months. Our case studies all started from data that already existed somewhere in the business.
What cleaning data costs, in the order you will meet it
“Clean the data” hides four different jobs with very different costs. Scoped to one use case rather than the whole business, this is what the work usually turns out to be.
- Getting access: days of work, weeks of waiting. The engineering is a connection string. The delay is legal review, a vendor who charges for API access, or an IT team with its own queue. Start this in week one regardless of what else you are doing, because it is the item most likely to set your timeline.
- Reconciling definitions: one to three weeks. Two systems disagree about what counts as an active customer. Somebody has to decide which is right, and that somebody is from your business, not from a vendor. This is the task companies consistently underestimate because it looks like a meeting rather than work.
- Fixing the records themselves: highly variable. If your sample of fifty came back at 90% accurate, you likely need spot fixes. At 60%, you need to understand why before touching anything, because the cause is usually upstream and will refill the errors within a quarter.
- Building the pipe that keeps it clean: one to two weeks. A one-off cleanup decays. Something has to keep the feed current, whether that is a scheduled job or a validation rule at the point of entry. Skipping this is why the second AI project in a company often meets the same data problems as the first.
Two of those four are your people rather than a vendor, which is the part proposals tend to leave out. When we quote a build, the access and definition work sits on the client side of the plan with named owners, and a project where nobody on the client side has time is a project we would rather delay than start.
When the answer is to fix the process, not add AI
Sometimes the data is bad because the process that produces it is bad, and then AI is the wrong purchase. Three signs. If the same information is typed into two systems by two people, the fix is an integration, not a model. If the field you need is free text because nobody made it a dropdown, the fix is a form change that costs a day. If the data is missing because the step that would capture it is skipped when people are busy, no model recovers what was never recorded.
We tell clients this in a meaningful share of audits, and it is the least commercial thing we do. It is also why the recommendation is worth something when it goes the other way: a firm that never says “don’t build this” is not assessing, it is quoting. If your project is stalling and you are not sure whether the cause is the data or the sequencing, the patterns behind failed AI projects covers the other common causes, most of which are also not technical.
- Data readiness is a property of a use case, not of your company. The same data is ready for retrieval and not ready for forecasting.
- Clean-then-build is the sequence that kills projects. Pick the use case that fits the data you have, ship it, and let it tell you which cleanup actually costs money.
- Five questions settle it: how many systems, who can approve access, structured or prose, machine-reachable, and does a sample of fifty records survive a manual check.
- Retrieval, classification, drafting and extraction tolerate messy data because a human is the last step. Forecasting, pricing and scoring do not, and produce confident wrong answers instead.
- When the same field is typed into two systems, or captured as free text, or skipped when people are busy, the honest fix is the process rather than a model.
An audit starts with the five questions above applied to two or three candidate use cases at once, which is what turns “is our data ready for AI” into a ranked list with costs attached. You can run it yourself with the questions on this page, or we can do it in a fixed fortnight and hand you the written result either way, including the version where the answer is to fix a form and spend nothing. Take your most promising use case: can you name the person who approves access to its data, and when did anyone last check fifty records by hand?
Book your AI audit

