The Wrong Report Card
Whenever we deploy something in an enterprise, sooner or later the CFO asks a very basic question: what is the ROI on this Agentic AI implementation?
Most of the time, we give them a number. The maths behind that number is also familiar. How many hours have been saved, how many headcounts have been avoided, what happened to AHT, how many days did the process reduce from, say, four days to one. These are all valid numbers and they matter. The problem is that these are largely the same numbers we used to measure RPA.
Those measures were built for something that follows rules. An RPA bot follows a predefined path, executes a task, and we measure how much faster or cheaper that task became. Agentic AI is different because the value is not only in doing the work. The real edge starts when the system can interpret context, make judgement calls, identify something unusual, and decide what should happen next.
That judgement is one of the reasons we are moving beyond traditional automation, but strangely, we are hardly measuring it.
The spreadsheet dashboards we built for automation are very good at telling us how much work got processed and how much time got saved. They are not very good at telling us whether the system made a better decision.
I was reading an excellent report by Eve Mitchell from SSON, The State of Intelligent Automation 2026. One of the findings was that most organisations running Agentic AI are still reporting less than 10% confirmed productivity improvement. The same report also raises the problem that traditional FTE and cycle-time metrics do not fully capture what these systems are actually delivering.
When you put those two things together, it starts to look like a different problem entirely.
Maybe the number isn’t low. Maybe the ruler is wrong.
Or perhaps more accurately, the ruler is incomplete.
Take a simple AP scenario. We deploy an agent to process invoices. Traditional RPA could also process invoices, extract the required fields, validate them, enter them into the ERP and route exceptions. So naturally, when we measure the new agent, we again calculate invoice processing time, AHT, number of invoices processed and hours saved.
Now imagine that while processing an invoice, the agent notices that the vendor bank account does not match the account historically used by that supplier. It flags the mismatch.
Maybe it is just a genuine bank account change. Maybe somebody entered incorrect master data. Maybe it is fraud. Changes to supplier payment details are a very real fraud vector, enough that the FBI specifically advises businesses to verify changes in account numbers and payment procedures. A tired human working on a Friday evening might have missed it, but the agent caught it.
Where does that show up in the ROI matrix?
Most likely, nowhere.
You cannot simply take the invoice value and say this is the amount the agent saved because you don’t know whether the payment would actually have gone through, whether somebody else would have caught it later, or whether there was any fraud at all. So there is no clean hard-dollar ROI line item for “the bad thing we avoided.”
And yet that one catch may be worth more than every hour the agent saved that month combined.
This is the part I think we are currently underselling, not overselling. We are so busy defending Agentic AI against the accusation that it has not delivered enough productivity that we are skipping the harder and probably more useful conversation: what should we actually be counting?
I don’t think the answer is to stop measuring ROI. If you cannot defend the spend with a number, you are unlikely to get the next budget cycle, and there is no point pretending otherwise. Hours saved, FTE avoidance, AHT, cycle time and processing cost still matter. What I am saying is that they should not be the complete report card for a system whose value increasingly comes from judgement.
A few things I have started recommending to track alongside the usual numbers.
How much faster is the decision being made?
AHT tells us how quickly somebody performs an activity. It doesn’t tell us how long the business waited for a decision.
A five-minute task may sit in somebody’s queue for eighteen hours. A procurement exception may wait for a manager. A customer complaint may move between three people because nobody has enough context to decide. A credit or pricing exception may sit for two days before someone acts on it.
If an agent reduces a five-minute activity to two minutes, the gain is three minutes. Useful, but not transformational.
If the same agent reduces a decision that normally sits unanswered for eighteen hours to twenty minutes, the operational effect can be far larger. Growth often slows not because the activity itself takes too long, but because decisions keep waiting.
So apart from AHT, we should probably start measuring decision latency: the time from “decision required” to “decision or action completed.”
That is a very different measure of productivity.
How well does the handoff work?
Another area we don’t measure well is the quality of the handoff.
An agent may complete its own work perfectly, but if it hands over an exception to a human without context, reasoning or a clear flag, it has not really saved much. It has simply moved the friction downstream.
If the person receiving the case has to reopen everything, understand what happened, check the data again and work out why the agent stopped, then the agent’s own processing time is almost irrelevant.
A good handoff should tell the human what happened, what the system checked, what looks unusual, why it stopped, what it recommends and what decision is actually required.
That can be measured too. How much time does the human spend after escalation? How many escalations are resolved without reopening the case? How many back-and-forth interactions are needed? How often does the human receive enough context to make the decision immediately?
The real gain is not only in how fast the agent completes its step. It is in whether the next person can act without starting again.
What did the system catch before it became a problem?
This is probably the hardest one to quantify and possibly the most valuable.
A near miss on compliance. A bad data entry stopped before it corrupts three downstream reports. A supplier fraud attempt flagged before payment. A duplicate payment caught. An unusual customer case identified before the wrong action is taken.
None of these necessarily gives you a neat rupee saving on the same day. They are costs avoided, and avoided costs are real costs even when accounting doesn’t have a clean column for them.
The mistake would be to manufacture a financial number for every catch. If the agent flags a ₹20 lakh invoice, it does not mean it has automatically saved ₹20 lakh.
But that does not mean we should record nothing.
We can start by counting the events. How many meaningful exceptions did the system catch? How many were genuine? What happened after they were flagged? What would historically have happened in similar cases? Over time, some of these can be connected to historical loss rates, error rates or expected exposure.
The first step is to stop treating “nothing bad happened” as equivalent to “nothing valuable happened.”
Are the humans doing better work now?
This is another area where the old report card can mislead us.
We tend to ask how many people the automation replaced. That is still a fair question, but it is not the only one.
Suppose a ten-person AP team is still a ten-person AP team after Agentic AI is deployed. On a traditional productivity dashboard, somebody may look at that and conclude that the technology has not delivered much because there was no headcount reduction.
But what if those ten people previously spent most of their time entering invoices, following up on routine items and checking standard exceptions, and now they spend much more of their time on supplier disputes, unusual cases, working-capital decisions and fixing root causes?
The headcount did not change. The quality of work did.
That should also be visible somewhere.
Instead of only asking how much human effort disappeared, it may be more useful to look at how human capacity shifted. How much time is now spent on repetitive processing versus exceptions, judgement, analysis and improvement?
The same team doing materially better work is also a productivity outcome, even if the FTE line remains unchanged.
This question came back to me after a recent NASSCOM NCR CXO Meet & Greet where we were discussing AI and productivity. One of the things I wrote after that discussion was that adoption is not the same as productivity. Enterprises can buy licences, run pilots, train people and deploy AI everywhere, but the more important question is what measurable difference all of that is actually making to the business.
The more I think about it, there is another question after that.
Are we measuring the right difference?
Because if AI is increasingly participating in judgement and decision-making, productivity cannot only mean how many minutes we removed from a task.
My follow-up meeting with one client actually ended better once we looked at the problem this way. We did not go back and say that the ROI number we had given them was wrong. It was accurate.
It was incomplete.
So we agreed to look at another set of numbers in the next quarter, not only what the agent had processed, but what it had caught, prevented, escalated better or helped the business decide faster.
That became a much better conversation.
Agentic AI should not be exempt from ROI scrutiny. If anything, the scrutiny should become better because enterprises are going to spend serious money on these systems.
But if the only report card we use was built to grade a rules-based bot, we should not be surprised when a decision-making system looks average on that report card even while it is quietly doing something valuable outside it.
Maybe we don’t need to throw away the RPA report card.
We just need to add the subjects that agents are actually being hired to perform.
If your leadership team is trying to work out where Agentic AI can genuinely improve operations, what is worth funding and how its value should be measured, our AI & Agentic Automation Leadership Workshop is designed around exactly these questions.




