Blog
How to measure AI ROI
The short answer
Pick one number that already exists in your business, measure it before you build, and change one thing. That is the entire method, and almost nobody follows it.
What usually happens instead is a spreadsheet assembled after the fact, multiplying an assumed time saving by a headcount by an hourly rate. That produces an impressive figure and convinces nobody who is paying attention.
Why the usual math is fiction
The standard model takes hours saved per person, multiplies by people, multiplies by cost per hour, and calls the product a return. Every term in that chain is an estimate, and the errors multiply rather than cancel.
Worse, saved hours are not recovered money unless something changed on the cost line or the output line. If the same team still works the same hours, the saving is real for them and invisible to finance, and finance is who approves the next phase.
A number your CFO can verify against a system they already trust beats a larger number you constructed.
Pick the number before you build
It should already be instrumented, already be reported, and already matter to someone senior. Conversion rate at a specific step. Cost per resolved ticket. Approval rate. Days to close. Error rate in a process with rework cost.
If it is not already measured, your first task is not AI, it is measurement, and pretending otherwise is how projects end up unable to prove anything. Write the number down, with its current value and the date, before a line of code exists.
The four shapes value takes
Revenue. More conversion, larger orders, better retention. Easiest to defend and hardest to attribute cleanly.
Cost. Fewer hours per unit, less rework, lower cost to serve. Verifiable if the cost was tracked before.
Risk. Fewer errors, better compliance, fewer incidents. Valued by what an incident costs, which finance can usually price.
Speed. Faster cycle time. Only real if the speed converts into one of the other three, so state which.
Baselines and the counterfactual
The hardest question you will be asked is whether the number would have moved anyway. Seasonality, a pricing change, a campaign, a competitor stumbling.
The clean answer is a holdout: run the new path for a portion of traffic and the old path for the rest, at the same time. If that is genuinely impossible, use a long enough pre period to show the trend, and be explicit about what you cannot rule out.
Stating the limitation yourself is far stronger than being caught by it. Executives discount numbers that arrive without caveats, correctly.
What to measure during the pilot
The outcome number, obviously. But also the operational ones that predict whether it survives: how often a human had to intervene, how often the system refused or escalated, and the quality trend on sampled traffic.
A pilot that hits the target while requiring constant intervention has not proven what people think it proved. Those signals come from LLM evals and observability, and without them you are reporting an outcome with no explanation attached.
The costs people forget
Denominators get built too small. Include the integration work, not just the build. Include running cost at production volume rather than pilot volume. Include the evaluation and monitoring you now maintain. Include the internal time, which is real money even when it does not appear on an invoice.
And include the maintenance you have signed up for, because a system in production is an obligation, not a purchase. The full breakdown is in what an AI agent costs.
When to kill it
Write the kill rule before you start, when nobody is invested. If the number has not moved by this date, we stop. Then honour it.
The failure mode here is not projects that fail, it is projects that quietly continue. A project with no kill rule cannot be killed, only starved, which wastes more than a clean stop and teaches the organisation that AI initiatives are permanent regardless of results.
Reporting it to a CFO
Lead with the metric, its before value, its after value, and the period. Then the cost, all of it. Then the caveat you cannot rule out. Then what you would do next and what it would cost.
Do not lead with the technology. A CFO does not need to know which model you used, and mentioning it first signals that the technology was the point rather than the outcome. That framing is also the difference between a partner and a vendor, which we covered in building in-house and in production AI, end to end.
If you need to justify an AI investment to someone who will challenge the math, thirty minutes will get you a defensible number. You leave with a price range and a clear next step.
Book a call Or send the details in writing