From Sanskriti Khandelwal | Product & Market Analysis
Negative ROI in AI Deployments: Two Forecast Errors That Never Cancel
On this page
Only 25% of AI initiatives have delivered the return their sponsors expected. In the same survey of 2,000 CEOs, 64% of chief executives admit they invest before they understand the value. Read together, the two findings describe a negative ROI deployment before it happens. The benefit was never specified, and the cost was never bounded.
Key takeaways
- Negative return and zero return are different events, and only one of them keeps billing. A pilot that returns nothing stops when the budget stops. A production deployment that returns less than it costs carries an inference bill, a supervision cost and a workflow that now depends on it.
- Both forecast errors point the same direction. IBM found 64% of CEOs invest before understanding the value, which inflates the benefit line. Forrester found fewer than one third of decision-makers can tie AI value to financial growth, which means the inflated line is never checked.
- The cost side grows after go-live, not before it. Human time spent fixing AI output was measured at $186 per employee per month by BetterUp Labs and Stanford, across 1,150 United States desk workers. That line appears in no business case I have seen.
- Cancellation is the cheap outcome, not the expensive one. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027. The projects that worry me are the ones nobody cancels because nobody measured them.
Zero return and negative return are not the same finding
Most public argument about AI returns uses one blunt category: the project that did not work. That category hides the distinction that matters to whoever signs the renewal.
Negative ROI in an AI deployment means the total twelve-month cost of running it exceeded the value it produced. It happens when the expected benefit is asserted rather than measured. It happens again when recurring costs are counted at pilot scale rather than production scale: inference, human review, integration work.
A deployment with zero return spent money and produced nothing measurable. A deployment with negative return spent money and produced something, and that something cost more to run and supervise than it saved. The second case is worse than the first, because it survives.
Pilots are cheap because they are small, short and closely watched by the people who wanted them. Production changes all three conditions at once, and every change points toward higher cost.
Volume rises, so the inference bill rises with it. Supervision moves from the enthusiastic project team to an operations team that never asked for the work. The evaluation harness built for the demo starts missing cases the demo never contained.
None of that shows up in month one. It shows up between month four and month nine. That is why twelve months is the honest window here, rather than the quarter after launch.
Error one: the benefit was never specified
The first error happens before anyone writes code. It is the more expensive of the two, because it cannot be corrected afterwards.
Two thirds of chief executives invest before they understand the value
The IBM Institute for Business Value surveyed 2,000 chief executives across 33 countries and 24 industries. Fieldwork ran from February to April 2025. The headline finding gets quoted often: only 25% of AI initiatives delivered their expected return, and just 16% scaled enterprise-wide.
The finding underneath it gets quoted far less. IBM reports that 64% of the same CEOs acknowledge that fear of falling behind drives investment in some technologies before they have a clear understanding of the value.
Read those two numbers together and the story changes. The 25% figure is usually presented as a delivery failure, as though the projects missed a fair target. But two thirds of buyers admit the target was set without understanding the value. That is a forecasting failure, not a delivery failure.
An unmeasured benefit defaults to zero at renewal
Here is the mechanism that turns a vague business case into a negative number. If nobody recorded the baseline before deployment, nobody can attribute the change afterwards. The finance team is then entitled to book the benefit at zero.
Forrester put a number on how common this is. In its 2026 predictions, published in October 2025, the firm found fewer than one third of decision-makers able to tie the value of AI to their organisation's financial growth. It predicted that enterprises would defer a quarter of planned AI spend into 2027 as a result.
That deferral prediction is the most useful line in the document. The correction arrives through budgets rather than through evidence. The evidence was never collected in a form a chief financial officer can audit.
The choice of benefit metric makes this worse. Most agent business cases are written in hours saved, and hours saved is the hardest benefit to defend. It converts to money only if the hours leave a cost line or move to revenue work.
Neither of those things happens automatically. A team that saves six hours a week and absorbs them into the same headcount has improved working life. It has not improved the profit and loss account. That is a legitimate outcome, and it is not a return.
Why measured productivity gains keep failing to appear in aggregate output is covered in the piece on the AI productivity paradox and GDP.
Error two: the cost line keeps growing after go-live
The second error is more forgivable. It involves costs that genuinely did not exist at approval time. It is still an error, and it is larger than most business cases allow for.
The clearest sign that this stopped being a small problem is organisational, not technical. The FinOps Foundation's 2026 survey covers 1,192 respondents representing more than $83 billion in annual cloud spend. It reports that 98% of practitioners now manage AI spend, up from 31% two years earlier.
A discipline does not go from 31% to 98% adoption over a rounding error. It moves when a line item starts breaking forecasts. The same survey names AI value management as the skillset its members are most actively trying to build.
Agentic workflows make this harder than chat-style usage. A single task fans out into a chain of model calls, retries and tool invocations. The mechanics of that bill are set out in the analysis of why cheaper tokens produce larger invoices. The forecasting problem itself is covered in the review of AI spend forecasting tools.
Supervision labour is the cost nobody puts in the model
BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full-time United States desk workers in September 2025. They found that 40% had received AI-generated work that looked finished but lacked substance. Resolving each incident took two hours on average.
Their cost estimate is $186 per employee per month. That annualises to roughly $9 million for a 10,000-person organisation. Treat it as directional: it rests on self-reported time and a salary assumption, not on measured throughput.
Directional is enough here. The cost is not zero, it recurs every month, and it lands on a different team from the one that bought the tool. Almost no business case contains a line for it.
A production agent embedded in a customer workflow has no natural exit. The cost of removing it grows with every month it runs.
The organisation ends up paying a recurring bill for a capability it now depends on. Nobody ever established that the capability pays for itself. Minimum commitments make this worse, which is the subject of the checklist of AI contract clauses a finance team should insist on.
Why the two errors never cancel each other out
Forecasting errors are normal and usually tolerable, because independent errors tend to offset one another. That reasoning does not apply here.
Both errors have the same cause: optimism about a technology the buyer has not yet run at scale. Optimism raises the benefit estimate and lowers the cost estimate in one act of judgement. The errors are correlated, so they add rather than offset.
A benefit overstated by 40% and a cost understated by 40% do not average out to roughly right.
| Mechanism | What it does to the forecast | Measured evidence |
|---|---|---|
| Investment approved before the value is understood | Raises the stated benefit | 64% of CEOs acknowledge it, IBM Institute for Business Value, 2,000 CEOs, 2025 |
| No baseline recorded, so no change can be attributed | Reduces the provable benefit toward zero | Fewer than one third of decision-makers can tie AI value to financial growth, Forrester, October 2025 |
| Inference priced at pilot volume, run at production volume | Understates recurring cost | 98% of FinOps practitioners now manage AI spend, up from 31%, FinOps Foundation, 1,192 respondents, 2026 |
| Human time spent checking and redoing AI output | Adds a recurring cost that was never modelled | $186 per employee per month, BetterUp Labs and Stanford, 1,150 US desk workers, 2025 |
These four figures come from four different surveys with four different populations. They cannot be summed into a single loss rate. The table shows direction, not magnitude.
Data access is a cost line, not a prerequisite
Missing data access is usually described as a blocker. In practice it rarely stops anything. The team routes around it, and the routing has a price.
When an agent cannot reach the system of record, the team builds an export, a copy or a narrower scope. Each of those adds engineering time and a synchronisation problem. Each also shrinks what the agent can do, which quietly reduces the benefit.
So the project ships. It ships with a smaller benefit than the business case assumed and a larger integration cost than the plan allowed. That is the negative ROI pattern arriving through a side door.
What the vendor surveys say, and how far to trust them
Cloudera commissioned Researchscape to survey 1,270 information technology leaders. All were at organisations with 1,000 or more employees, across the Americas, Europe and Asia Pacific, between January and March 2026. It reported that nearly 80% of enterprises say their AI and data initiatives are constrained by limited data access across environments.
That is vendor-commissioned research from a company that sells data platforms, so it is not a load-bearing figure here. It is worth reporting because the same survey contains a finding that cuts against the sponsor's interest. In it, 84% of respondents felt confident in the accuracy and completeness of their data.
Confidence in data quality is high, and access to that data is constrained. A business case written under the first belief meets the second reality in month three.
What the cancellation data actually tells you
Cancellation figures get used as evidence of an AI reckoning. I read them close to the opposite way. The difference matters for how you run your own portfolio.
Gartner published its cancellation forecast on 25 June 2025, naming escalating costs, unclear business value and inadequate risk controls as the causes. Anushree Verma, a senior director analyst at the firm, described most current agentic projects as early-stage experiments driven by hype and often misapplied.
The same release estimates that of the thousands of vendors claiming agentic capability, only around 130 offer something genuine. That figure belongs in every procurement conversation. A deployment built on a repackaged workflow tool inherits the cost of an agent and the capability of a script.
Abandonment rose sharply, and that is a healthy signal
S&P Global Market Intelligence surveyed more than 1,000 respondents across North America and Europe. It found the share of businesses scrapping most of their AI initiatives rising to 42% from 17% a year earlier. The average organisation scrapped 46% of its proofs of concept before production.
A rising abandonment rate is the market learning to measure. An organisation that cancels 46% of its experiments is running a portfolio. One that cancels almost nothing is either extraordinarily good at selection or not checking at all.
The pilot-stage version of this question is handled separately in the breakdown of six agent pilot failure modes. That post covers which failure modes stop projects before production, and why the widely quoted 88% figure has no source. This post picks up where it stops, at the deployments that did ship.
Forrester reported in June 2026 that three quarters of enterprise leaders say they are adopting agentic AI. Only a small minority run it in meaningful production beyond chatbot-style deployments. The same analysis found more than half of enterprises reporting agent sprawl, even after adopting a recognised governance framework.
Sprawl is a cost word. It means agents running without a named owner, without a budget line and without anyone computing what they return. That is the precise condition under which a negative number goes unnoticed for a year.
Where this argument is weakest
Three things about the case above deserve to be said out loud. The first concerns a number in the brief for this post.
The 22% figure could not be verified
A figure circulating widely holds that 22% of agent deployments report negative return at twelve months. The claimed root-cause split is 41% unclear success criteria, 33% insufficient tool or data access and 26% evaluation drift. It is usually attributed to Forrester.
I could not trace it. Every restatement found during this research sat on an aggregator page. None linked to a Forrester report, press release or briefing with a stated sample. Forrester's own published material on agentic AI in 2026 contains no such figure. So this post does not use it, and neither should you until somebody produces the source.
Every survey here has a survivorship problem
Each of these datasets asks people inside organisations about outcomes those same people are partly responsible for. Self-reported return is a claim, not an audit. The direction of the bias is not obvious.
Respondents may overstate returns to protect a decision they championed. They may equally understate them, because unmeasured benefits are invisible to the person filling in a form. No published dataset found during this research resolves that.
The case that negative return is the correct price
The strongest argument against this framing is simple. A twelve-month accounting window is the wrong instrument for a capability build. Learning what an agent cannot do is worth paying for, and that lesson arrives as a cost.
I accept that case for a bounded number of experiments with a stated learning goal and a stated stop date. I do not accept it for a production deployment running for a year. At that point the organisation is not learning, it is subsidising. IBM's own respondents expect the subsidy to end: 85% told the firm they expect positive returns from scaled efficiency investments by 2027.
A twelve-month ledger you can actually run
The remedy for both forecast errors is the same short document. Write the ledger before approval. Run it again at month twelve against exactly the same lines.
| Line | Usually in the business case? | Where the real number lives |
|---|---|---|
| Licence or platform fee | Yes | The contract, including any minimum commitment |
| Inference at production volume | Rarely, and usually at pilot rates | Provider usage console, measured over a full month at real volume |
| Human review and rework time | No | Ask the receiving team to log it for two weeks |
| Integration and data access engineering | Partly, and usually as a one-off | Engineering time sheets across the first two quarters |
| Incident response and rollback | No | Operations ticket volume tagged to the deployment |
| Cost of removing it later | No | Estimate the workflow rebuild before you depend on the workflow |
| Benefit, against a recorded baseline | Asserted, not measured | The metric value from the month before launch, written down |
The last row decides everything else. If it is empty at approval, the deployment cannot produce a positive result later. There will be nothing to compare against. The related economics are set out in the analysis of the AI unit economics margin trap.
One more discipline costs nothing. Write down the number below which you will stop, and the date you will check it, in the document that approves the spend. A stop condition agreed in advance is the only version of this decision that does not become political later.
Where measurable return has shown up, and what those cases share, is examined in the piece on who is actually making money from generative AI.
Frequently asked questions
What is negative ROI in AI?
Negative ROI in AI means the total cost of running a deployment over a period exceeded the value it produced in the same period. It differs from zero return, where a pilot produced nothing measurable and then stopped. A negative-return deployment keeps operating, so it keeps generating inference charges, supervision time and integration work while the benefit stays unproven.
What percentage of AI projects have negative ROI?
No published survey measures that directly, and figures circulating with a specific share should be checked against a primary source before use. What is measured is adjacent. IBM found 25% of AI initiatives delivered their expected return across 2,000 CEOs surveyed in 2025. S&P Global found 42% of businesses scrapping most AI initiatives, up from 17% the year before.
Why do AI agent deployments lose money?
Two forecast errors compound. The benefit is set before the buyer understands the value, which IBM found 64% of CEOs admit to. The cost is estimated at pilot scale and then runs at production scale, adding inference volume, human review time and integration work. Because both errors come from the same optimism, they add together rather than cancelling out.
How do you measure ROI on an AI agent?
Record the baseline value of the target metric in the month before launch, with the period it covers. Track seven cost lines rather than one: licence, inference at real volume, human review time, integration engineering, incident response, removal cost and any minimum contract commitment. Compare at twelve months, not at the quarter after launch, because most cost growth appears between months four and nine.
When should you cancel an AI project?
Cancel when the twelve-month ledger is negative and no line has a credible path to changing. Set that stop number and its review date in the approval document, before anyone is invested in the outcome. Gartner expects over 40% of agentic projects to be cancelled by the end of 2027, so treat cancellation as a normal portfolio outcome rather than a failure event.
Is AI cost overrun the main reason projects fail?
Cost is one of three causes Gartner names, alongside unclear business value and inadequate risk controls. In practice the benefit side does more damage, because an unmeasured benefit defaults to zero when finance reviews it. Forrester found fewer than one third of decision-makers able to tie AI value to financial growth, which makes almost any cost figure look unjustified by comparison.
Where to start this week
Pick your single largest AI deployment and answer one question about it: what was the number before you turned it on. Not the estimate, the recorded value of the metric it was meant to move, in the month before launch.
If that number exists, run the seven-line ledger above against it this quarter. If it does not exist, you have found the finding. Record the baseline now, and treat the last twelve months as unmeasurable rather than as a success.
Then put one date in the calendar: twelve months from launch, or three months from today if launch was long ago. Write the stop number next to it.
Related analysis
This piece covers deployments that shipped. For the stage before that, read the six failure modes that stop agent pilots reaching production. For the cost line that grows fastest, read why cheaper tokens produce larger bills.
References
- IBM Institute for Business Value, CEOs double down on AI while navigating enterprise hurdles, 6 May 2025. 2,000 CEOs, 33 countries, 24 industries, fieldwork February to April 2025. Used for the 25%, 16%, 64% and 85% figures.
- Forrester, 2026 Technology and Security Predictions, 28 October 2025. Used for the deferred AI spend prediction and the share of decision-makers who can tie AI value to financial growth.
- Forrester, The state of agentic AI in 2026, 3 June 2026. Used for adoption against production figures and the agent sprawl finding.
- Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation forecast, the causes, the January 2025 poll of 3,412 attendees and the vendor count.
- CIO Dive, AI project failure rates are on the rise, 14 March 2025, reporting S&P Global Market Intelligence Voice of the Enterprise research, more than 1,000 respondents in North America and Europe. Used for the 42% and 46% figures.
- BetterUp Labs and Stanford Social Media Lab, Workslop research, survey of 1,150 full-time United States desk workers, September 2025, first published in Harvard Business Review. Used for the $186 monthly figure and the two-hour resolution time.
- FinOps Foundation and Linux Foundation, State of FinOps 2026 survey, 19 February 2026. 1,192 respondents representing more than $83 billion in annual cloud spend. Used for the 98% and 31% figures.
- Cloudera, Data Readiness Index, 14 April 2026, fieldwork by Researchscape, 1,270 IT leaders, 22 January to 3 March 2026. Vendor-commissioned. Used for the data access and data confidence figures, both labelled as such in the text.
The weakest thing about this source base: every figure here is self-reported by people with a stake in the outcome, and none of it is an audited financial result. Two of the eight sources are forecasts rather than measurements, and one is vendor-commissioned. The waterfall figure is illustrative and contains no measured amounts.
Related reading