Measuring ROI on AI Automation: a 90-Day Framework
A practical framework for proving the business case for AI automation: four value types, the costs people forget, and what to measure.
The short answer
Claim value in four labelled types, in order of defensibility: cost avoided, revenue enabled, risk reduced, capacity created. A business case that mixes all four into one number is not measurable. Baseline the process for at least two weeks before anything changes, then run 90 days: measure, build and shadow, go live with human oversight, then compare against baseline and decide. Stopping at day 90 with a clear answer is a legitimate and cheap outcome.
Key points
- Saved hours only become money if the time is redeployed or headcount genuinely changes. Otherwise report it as capacity created.
- First-year running cost is a meaningful fraction of build cost, not a rounding error. A case that omits it is a proposal, not an analysis.
- Two weeks of imperfect measurement beats a confident estimate.
- Shadow mode, running the system alongside humans without acting on its output, is the stage most teams skip and most regret skipping.
- Report one primary metric, total cost including running cost, the value type claimed, and one honest sentence about what has not worked.
Most AI business cases are written to get approval rather than to find the truth. They count the hours saved and ignore the cost of running the thing. Then nobody revisits the numbers, so the organisation never learns which projects were actually worth doing.
Here is a version designed to be honest.
The four types of value, in order of defensibility
Cost avoided. Work that no longer needs doing. The easiest to measure and the easiest to overstate, because saved hours only become money if the time is redeployed or headcount genuinely changes.
Revenue enabled. Faster response times, more leads worked, fewer opportunities dropped. Harder to attribute, usually larger. Improvements in speed to lead are the most reliable example.
Risk reduced. Fewer errors, better compliance, fewer things falling through gaps. Rarely quantified, occasionally the entire justification.
Capacity created. The ability to handle more volume without proportional cost. The hardest to prove and the most valuable over several years.
Claim value in this order and label which type you are claiming. A business case mixing all four into one number is not measurable.
The costs people forget
Build cost is the visible number. The rest: model and infrastructure usage, integration and licence costs, internal time during discovery and testing, the change management effort, ongoing monitoring and iteration, and the periodic fix when an upstream system changes.
A reasonable planning assumption is that first-year running cost is a meaningful fraction of build cost, not a rounding error. Any business case that omits year one running cost is a proposal, not an analysis.
Baseline before you build
If you do not know the current number, you cannot prove improvement, and you will end up arguing about whether it feels better.
Measure for at least two weeks before anything changes: volume per period, minutes per instance, error and rework rate, cycle time end to end, and the human cost of that time. Two weeks of imperfect measurement beats a confident estimate.
The 90-day structure
Days 1 to 14: baseline. Measure the current process, agree the target metric, write down what success looks like numerically.
Days 15 to 45: build and shadow. Build, then run the system alongside the humans without acting on its outputs. Compare quality. This is the stage most teams skip and most regret skipping.
Days 46 to 75: live with oversight. The system does the work, a human reviews before anything final. Track error rate and how often review changes the output.
Days 76 to 90: measure and decide. Compare against baseline. Decide to expand, adjust or stop.
Stopping is a legitimate outcome. A project that ends at day 90 with a clear answer cost you a quarter. One that limps along unmeasured for two years costs considerably more.
What to report to leadership
One primary metric with its baseline and current value. Total cost including running cost. Type of value claimed. One honest sentence about what has not worked. That last item is what makes the rest of the report believable.
Choosing what to measure first
The framework is only as good as the process you point it at. If you are still deciding, the first five processes companies in India and the UAE automate are the ones that baseline most easily. And if an external partner is doing the work, a willingness to be measured this way is one of the clearest signals of a serious agency.
SolvTree scopes AI work against a baseline and reviews it against the same numbers at the end. If you want the honest version of your business case, talk to us.
Frequently asked questions
- What is a reasonable payback period for AI automation?
- For a scoped operational automation, aim to justify it within a year including running costs. Longer horizons can be defensible for platform work, but they should be argued explicitly rather than assumed.
- How do we value time saved?
- Only count it if the time genuinely goes somewhere else you can name. Otherwise report it as capacity created and be honest that it is not yet cash.
- What if the numbers are unimpressive?
- Then you have learned something cheaply. That is the point of measuring at ninety days rather than at year three.