AI Strategy
Where AI actually creates ROI
Most AI programmes produce impressive demos and disappointing returns. The difference is rarely the model — it is which process the work was aimed at.
There is a familiar pattern in enterprise AI. A pilot is commissioned. It works. Everyone is impressed. Eighteen months later it has not been rolled out, nobody can say what it saved, and the budget has quietly moved elsewhere.
The instinct is to blame execution — change management, data quality, integration. Sometimes that is right. More often the pilot was pointed at a process where AI was never going to pay, and the demonstration succeeded precisely because it was measuring the wrong thing.
Why demos mislead
A demo proves the model can do the task. That is not the question.
The question is whether the model can do the task at sufficient accuracy, at acceptable cost, on the real distribution of inputs, inside a workflow someone will actually adopt, at a volume that matters. A demo answers the first clause and is silent on the rest — and the rest is where the return lives.
The gap shows up in four places, reliably:
The real input distribution is worse than the demo set. Demos are built on clean examples. Production brings scanned faxes, three languages in one document, and the edge cases that make up a third of the volume and all of the difficulty.
The accuracy threshold is a cliff, not a slope. For many processes there is a level below which output requires full human review — at which point you have added cost, not removed it. Going from 80% to 92% may be worth nothing at all if the threshold sits at 95%.
Cost per inference is a real operating line. At demo volumes it rounds to zero. At production volumes it is a line item that has to be smaller than the thing it replaces, and frequently is not.
The workflow does not change itself. A model that produces a good answer into a system nobody uses has produced nothing. This is the most common cause of stalled rollout and the least discussed in the business case.
The four places it consistently pays
Across the work I have seen, the returns cluster. Not around industries — around process characteristics.
1. High-volume, low-value-per-unit judgement
Work that is individually trivial and collectively enormous. Categorising inbound requests, matching records, extracting fields from documents, routing tickets, first-pass content moderation.
These pay because volume amortises the build, the accuracy threshold is usually forgiving, and the alternative — humans doing repetitive judgement work — is expensive and gets worse under fatigue. This is the least glamorous category on the list and the most reliable.
2. Work currently not done at all
The strongest cases are frequently not automation. They are things a business would like to do and cannot afford to.
Reviewing every contract rather than a sample. Inspecting the whole road network rather than the inspected fraction. Giving every customer the analysis currently reserved for the top tier. There is no incumbent cost to compare against, so the business case is pure upside — and the accuracy bar is set against nothing, which is a much easier bar than against a skilled human.
3. Unstructured data that gates a decision
Organisations hold enormous quantities of text, images and audio that never enter a decision because reading it does not scale. Call recordings, field notes, inspection photographs, support transcripts, listing descriptions.
The value is not the extraction. It is that a decision previously made on partial information can now be made on complete information. Look for places where someone samples because they cannot read everything.
4. Latency-critical judgement
Decisions where being right in two seconds is worth substantially more than being right in two hours. Fraud interdiction at the point of transaction, dynamic pricing, real-time routing, next-best-action while the customer is still on the line.
Here AI is not replacing a human doing the same thing — it is replacing a decision that was previously made too late to matter, or made by a rules engine that could not carry the complexity.
Where it consistently does not pay
Equally worth naming.
Low-volume, high-stakes judgement. If a decision happens forty times a year and each one is consequential, a human will make it. The volume never amortises the build and the accuracy bar is brutal.
Processes that are broken for organisational reasons. If the approval chain has eleven steps because of a governance dispute, an AI that accelerates step four saves nothing. The constraint is not throughput.
Anything where the incumbent process is already cheap. A well-built rules engine handling 95% of cases at near-zero marginal cost is a hard thing to beat. It is unfashionable to say so, and it is frequently true.
Work where the output requires full verification anyway. If a human must check every result against the source to the same depth, the model has added a step.
How to evaluate a case before you fund it
Five questions, in order. If the first three do not have numbers attached, the business case is not finished.
- What is the volume, and what is the current fully-loaded cost per unit? Without this there is no baseline and therefore no possible return.
- What accuracy is required for the output to be usable without full review? Find the cliff before you build, not after.
- What will inference cost at full volume? Model it at the real rate, including retries and the long tail of hard inputs.
- Who changes their behaviour, and have they agreed to? Name the team. If nobody’s daily work changes, nothing has been saved.
- How will you know it worked? A metric that already exists, already has a baseline, and is already reported. If the measurement has to be invented alongside the project, it will not survive the first review.
The underlying point
AI does not create return. Applying it to a process where the economics were already broken creates return. The technology is genuinely remarkable and that is exactly what makes it easy to point at problems that did not need it.
The organisations getting real value are not the ones with the best models. They are the ones that did the unglamorous work of finding out where their money was actually going first.

Vidhu Saxena
Founder & Principal Product Consultant, AithozPM
Twenty years building and running digital products across PropTech, FinTech, marketplaces and enterprise software in India, the GCC and Europe.
More about AithozPM