Measuring AI ROI Without Fooling Yourself
The baseline cannot be collected after launch, which makes two weeks of measurement before the build the whole discipline. The four things to record, why hours saved is not money until a named line in the accounts moves, the metrics that survive a board, and the kill criteria to write down first.

The number that decides this cannot be collected later.
Measuring AI ROI is almost entirely a question of what you wrote down before the build started. Afterwards the baseline is gone. People reconstruct it from memory, memory is generous in whichever direction the project needs, and the resulting figure convinces nobody who was not already convinced. Two weeks of measurement before anyone commits code is the difference between a defensible result and a slide that gets politely accepted and quietly ignored. And when you do have the number, expect the honest version to be smaller and more specific than the one in the proposal, which is exactly why it survives scrutiny.
The AI ROI baseline is four things, measured over two weeks
An AI ROI baseline is not a workshop estimate. It is actual observation of the process as performed, for long enough to catch a normal range rather than a good week.
- Volume. How many of these happen a week, and how that varies. Almost every business overestimates the volume of the process they want automated, sometimes by a factor of three, and volume is what the entire return multiplies against.
- Time per unit, end to end. Not the time somebody spends typing, the elapsed time from arrival to done, including the two days it sat in a queue. The queue is usually where the value is, and it is the part nobody thinks to time.
- Error and rework rate. How often it comes back. Measure it now, because after launch every error will be attributed to the new system and nobody will be able to say whether it is better or worse than before.
- Who does it, and what else they would do. Names and roles. This is the line that converts time into money later, and without it you will be left with hours that cannot be priced.
Write down, on the same day, every other change happening to this process in the next six months. A new hire, a system migration, a pricing change, a seasonal peak. You will not remember them when the results come in, and they are the first thing a skeptical CFO will ask about. This measurement is the first stage of any competent implementation and the reason our audit produces a number before it produces a specification.
Why hours saved rarely survives a finance review
Hours saved is the standard AI benefit and the easiest to challenge, because saved hours are not money until something changes. Twenty people each saving two hours a week is forty hours, and it is forty hours of nothing unless a headcount changes, unless the freed time goes to work that is billed, or unless a queue that was capping revenue stops capping it. In most companies none of those three happens by default. The time is absorbed, the work expands, and a year later the finance director cannot find the saving in any account.
That is not an argument for skipping it. It is an argument for finishing the sentence. Convert hours saved into one of three things and it becomes real: a headcount you did not need to add for a role you can name, billable hours that were previously administrative in a business that bills, or throughput you could not previously reach because the process was the constraint. If the honest answer is that none of the three applies, say so and argue the project on service quality or risk instead. Those are legitimate cases and they are far more persuasive than an unconvertible hours figure.
A quick test before the number goes in a deck. Ask which line in the accounts changes, and by when. If nobody can name the line, the benefit is real to the team and invisible to the business, and the project needs a different justification rather than a bigger hours estimate.
AI ROI metrics that hold up, and the ones that flatter
Metrics that survive scrutiny share one property: they were being measured before the project existed, usually by somebody with no stake in it. Cost per unit of work. Elapsed cycle time from arrival to resolution. Rework rate. The share of total volume completed without a human touching it. Throughput per person where headcount is genuinely fixed. Each of those has a before, an after, and an owner outside the project.
The flattering ones all measure the system rather than the business. Number of people who tried it. Queries answered, which counts activity and not resolution. Model accuracy in isolation, which can be excellent while the workflow around it fails. Time to first response, without whether the thing was ever actually resolved. Any of these can be true and rising while the process is no better than it was, and a metric that cannot get worse is not a measurement.
The Biscoito.ai veterinary assistant is a fair example of the difference. The number it reports is that 60 to 70 per cent of owner enquiries are resolved in under five seconds, around the clock. Resolved is doing the work in that sentence: it is a share of total volume that reached an outcome, not a count of conversations started, and the clinic’s original problem was weekend bookings lost because nobody answered the phone, which is a revenue line rather than an hours line. Our other case studies are written against the same standard.
On attribution, three practical options, in descending order of rigour and ascending order of convenience. Roll out in stages by team or region and compare the ones that have it against the ones that do not, which is the cheapest genuine experiment available and costs only sequencing. Hold out a group deliberately for six weeks. Or compare against the same period last year and publish, alongside the result, the list of other changes you wrote down at the start. The third is weak and it is honest, which beats a strong claim that falls over under one question.
What a board wants, and when to kill the project
A board wants four lines and no dashboard: what we spent including the run cost, what changed in a number that existed before we started, what we are not yet able to attribute, and what we would do differently. The fourth line is what buys you credibility for the next proposal, and the third is what stops a skeptic dismantling the first two in public. Turning that into a document somebody will approve is its own exercise, covered in the AI business case your board will approve.
Killing a project is the part nobody plans and it needs its own criteria, set at the start. Three signals are usually decisive. Adoption sits below half of the eligible volume twelve weeks after launch with no process change in progress, which means the system is being routed around rather than used. The target metric moved and the business line it was supposed to drive did not, which means you improved something that was not the constraint. Or the baseline was never taken and cannot be reconstructed, in which case you are not going to be able to prove anything, and the honest options are to restart with measurement or to stop.
Set those thresholds before launch and write them in the same document as the baseline. Deciding what counts as failure while looking at results is not a decision, it is a negotiation with yourself.
Two cases where ROI is the wrong frame
Small, cheap and obviously correct. A €4,000 automation that removes a task everybody agrees is wasteful does not need a measurement programme costing a quarter of the build. Take the baseline anyway, because it is two hours of work, and skip the rest.
The driver is risk or obligation. A system built because a regulator, a client contract or an insurer requires something is not an investment case, and forcing it into one produces a fictional number that weakens a real argument. Say what the requirement is and what non-compliance would cost, then measure whether the obligation is being met rather than whether it paid back.
- The baseline cannot be collected after launch. Two weeks of observation before anyone writes code decides whether the result is defensible.
- Record four things: volume, elapsed time per unit including queue time, rework rate, and who does it. Then list every other change coming in the next six months.
- Hours saved is not money until a headcount changes, the time becomes billable, or a constraint on throughput lifts. Name the line in the accounts that moves, or argue the project on quality or risk instead.
- The AI ROI metrics that hold up were already being measured by somebody with no stake in the project. Users who tried it, queries answered and model accuracy in isolation are not measurements.
- Write the kill criteria next to the baseline. Deciding what counts as failure while looking at the results is a negotiation, not a decision.
The cheapest way to measure AI ROI is to spend two weeks measuring the process before anybody builds anything, which is what the first half of our audit does. You leave with volume, cycle time, rework rate and a cost per unit, in a document you own, and that document is just as useful for deciding not to build as it is for justifying the build afterwards. A fair share of the time it argues for the smaller project. One question to test where you stand: if the system you are planning launched next month, could you say today what number it would be compared against?
Book your AI audit

