From Sanskriti Khandelwal | Product & Market Analysis

Negative ROI in AI Deployments: Two Forecast Errors That Never Cancel

On this page

Only 25% of AI initiatives have delivered the return their sponsors expected. In the same survey of 2,000 CEOs, 64% of chief executives admit they invest before they understand the value. Read together, the two findings describe a negative ROI deployment before it happens. The benefit was never specified, and the cost was never bounded.

Key takeaways

  • Negative return and zero return are different events, and only one of them keeps billing. A pilot that returns nothing stops when the budget stops. A production deployment that returns less than it costs carries an inference bill, a supervision cost and a workflow that now depends on it.
  • Both forecast errors point the same direction. IBM found 64% of CEOs invest before understanding the value, which inflates the benefit line. Forrester found fewer than one third of decision-makers can tie AI value to financial growth, which means the inflated line is never checked.
  • The cost side grows after go-live, not before it. Human time spent fixing AI output was measured at $186 per employee per month by BetterUp Labs and Stanford, across 1,150 United States desk workers. That line appears in no business case I have seen.
  • Cancellation is the cheap outcome, not the expensive one. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027. The projects that worry me are the ones nobody cancels because nobody measured them.
25%Share of AI initiatives that delivered their expected return. Just 16% scaled across the enterprise. Source: IBM Institute for Business Value, May 2025.
1 in 3Fewer than one in three decision-makers can tie AI value to financial growth. Source: Forrester Predictions 2026, October 2025.
$186Monthly cost per employee of resolving low-quality AI output. Source: BetterUp Labs and Stanford Social Media Lab, September 2025.

Zero return and negative return are not the same finding

Most public argument about AI returns uses one blunt category: the project that did not work. That category hides the distinction that matters to whoever signs the renewal.

Negative ROI in an AI deployment means the total twelve-month cost of running it exceeded the value it produced. It happens when the expected benefit is asserted rather than measured. It happens again when recurring costs are counted at pilot scale rather than production scale: inference, human review, integration work.

A deployment with zero return spent money and produced nothing measurable. A deployment with negative return spent money and produced something, and that something cost more to run and supervise than it saved. The second case is worse than the first, because it survives.

Pilots are cheap because they are small, short and closely watched by the people who wanted them. Production changes all three conditions at once, and every change points toward higher cost.

Volume rises, so the inference bill rises with it. Supervision moves from the enthusiastic project team to an operations team that never asked for the work. The evaluation harness built for the demo starts missing cases the demo never contained.

None of that shows up in month one. It shows up between month four and month nine. That is why twelve months is the honest window here, rather than the quarter after launch.

Error one: the benefit was never specified

The first error happens before anyone writes code. It is the more expensive of the two, because it cannot be corrected afterwards.

Two thirds of chief executives invest before they understand the value

The IBM Institute for Business Value surveyed 2,000 chief executives across 33 countries and 24 industries. Fieldwork ran from February to April 2025. The headline finding gets quoted often: only 25% of AI initiatives delivered their expected return, and just 16% scaled enterprise-wide.

The finding underneath it gets quoted far less. IBM reports that 64% of the same CEOs acknowledge that fear of falling behind drives investment in some technologies before they have a clear understanding of the value.

Read those two numbers together and the story changes. The 25% figure is usually presented as a delivery failure, as though the projects missed a fair target. But two thirds of buyers admit the target was set without understanding the value. That is a forecasting failure, not a delivery failure.

An unmeasured benefit defaults to zero at renewal

Here is the mechanism that turns a vague business case into a negative number. If nobody recorded the baseline before deployment, nobody can attribute the change afterwards. The finance team is then entitled to book the benefit at zero.

Forrester put a number on how common this is. In its 2026 predictions, published in October 2025, the firm found fewer than one third of decision-makers able to tie the value of AI to their organisation's financial growth. It predicted that enterprises would defer a quarter of planned AI spend into 2027 as a result.

That deferral prediction is the most useful line in the document. The correction arrives through budgets rather than through evidence. The evidence was never collected in a form a chief financial officer can audit.

The choice of benefit metric makes this worse. Most agent business cases are written in hours saved, and hours saved is the hardest benefit to defend. It converts to money only if the hours leave a cost line or move to revenue work.

Neither of those things happens automatically. A team that saves six hours a week and absorbs them into the same headcount has improved working life. It has not improved the profit and loss account. That is a legitimate outcome, and it is not a return.

Why measured productivity gains keep failing to appear in aggregate output is covered in the piece on the AI productivity paradox and GDP.

Error two: the cost line keeps growing after go-live

The second error is more forgivable. It involves costs that genuinely did not exist at approval time. It is still an error, and it is larger than most business cases allow for.

The clearest sign that this stopped being a small problem is organisational, not technical. The FinOps Foundation's 2026 survey covers 1,192 respondents representing more than $83 billion in annual cloud spend. It reports that 98% of practitioners now manage AI spend, up from 31% two years earlier.

A discipline does not go from 31% to 98% adoption over a rounding error. It moves when a line item starts breaking forecasts. The same survey names AI value management as the skillset its members are most actively trying to build.

Agentic workflows make this harder than chat-style usage. A single task fans out into a chain of model calls, retries and tool invocations. The mechanics of that bill are set out in the analysis of why cheaper tokens produce larger invoices. The forecasting problem itself is covered in the review of AI spend forecasting tools.

Supervision labour is the cost nobody puts in the model

BetterUp Labs and the Stanford Social Media Lab surveyed 1,150 full-time United States desk workers in September 2025. They found that 40% had received AI-generated work that looked finished but lacked substance. Resolving each incident took two hours on average.

Their cost estimate is $186 per employee per month. That annualises to roughly $9 million for a 10,000-person organisation. Treat it as directional: it rests on self-reported time and a salary assumption, not on measured throughput.

Directional is enough here. The cost is not zero, it recurs every month, and it lands on a different team from the one that bought the tool. Almost no business case contains a line for it.

A production agent embedded in a customer workflow has no natural exit. The cost of removing it grows with every month it runs.

The organisation ends up paying a recurring bill for a capability it now depends on. Nobody ever established that the capability pays for itself. Minimum commitments make this worse, which is the subject of the checklist of AI contract clauses a finance team should insist on.

How a twelve-month deployment lands below zero Illustrative worked example. The shape is the argument, not the amounts. Forecast benefit $400k Benefit not attributable -$180k Inference above plan -$90k Supervision and rework -$85k Access and integration -$70k Net… month 12 -$25k Only the first bar was in the business case. The four deductions arrived after go-live.
This chart shows a relationship, not measured data. No single deduction is fatal on its own. Four ordinary ones together are.

Why the two errors never cancel each other out

Forecasting errors are normal and usually tolerable, because independent errors tend to offset one another. That reasoning does not apply here.

Both errors have the same cause: optimism about a technology the buyer has not yet run at scale. Optimism raises the benefit estimate and lowers the cost estimate in one act of judgement. The errors are correlated, so they add rather than offset.

A benefit overstated by 40% and a cost understated by 40% do not average out to roughly right.

Two forecast errors, four mechanisms, one direction
MechanismWhat it does to the forecastMeasured evidence
Investment approved before the value is understoodRaises the stated benefit64% of CEOs acknowledge it, IBM Institute for Business Value, 2,000 CEOs, 2025
No baseline recorded, so no change can be attributedReduces the provable benefit toward zeroFewer than one third of decision-makers can tie AI value to financial growth, Forrester, October 2025
Inference priced at pilot volume, run at production volumeUnderstates recurring cost98% of FinOps practitioners now manage AI spend, up from 31%, FinOps Foundation, 1,192 respondents, 2026
Human time spent checking and redoing AI outputAdds a recurring cost that was never modelled$186 per employee per month, BetterUp Labs and Stanford, 1,150 US desk workers, 2025

These four figures come from four different surveys with four different populations. They cannot be summed into a single loss rate. The table shows direction, not magnitude.

Five measured shares, three different denominators Read the right-hand column before comparing any two bars. They do not measure the same population. Delivered expected ROI 25% of AI initiatives, IBM 2025 Scaled enterprise-wide 16% of AI initiatives, IBM 2025 Invest before value is clear 64% of CEOs, IBM 2025 Can tie AI to financial growth under 33% of decision-makers, Forrester 2025 Predicted cancelled by 2027 over 40% of agentic projects, Gartner 2025 The pale bar is a forecast, not a measurement. The other four are survey results.
The 64% bar is the one to sit with. It is a self-reported admission from the people approving the spend. That makes it harder to argue away than the outcome figures.

Data access is a cost line, not a prerequisite

Missing data access is usually described as a blocker. In practice it rarely stops anything. The team routes around it, and the routing has a price.

When an agent cannot reach the system of record, the team builds an export, a copy or a narrower scope. Each of those adds engineering time and a synchronisation problem. Each also shrinks what the agent can do, which quietly reduces the benefit.

So the project ships. It ships with a smaller benefit than the business case assumed and a larger integration cost than the plan allowed. That is the negative ROI pattern arriving through a side door.

What the vendor surveys say, and how far to trust them

Cloudera commissioned Researchscape to survey 1,270 information technology leaders. All were at organisations with 1,000 or more employees, across the Americas, Europe and Asia Pacific, between January and March 2026. It reported that nearly 80% of enterprises say their AI and data initiatives are constrained by limited data access across environments.

That is vendor-commissioned research from a company that sells data platforms, so it is not a load-bearing figure here. It is worth reporting because the same survey contains a finding that cuts against the sponsor's interest. In it, 84% of respondents felt confident in the accuracy and completeness of their data.

Confidence in data quality is high, and access to that data is constrained. A business case written under the first belief meets the second reality in month three.

What the cancellation data actually tells you

Cancellation figures get used as evidence of an AI reckoning. I read them close to the opposite way. The difference matters for how you run your own portfolio.

Gartner published its cancellation forecast on 25 June 2025, naming escalating costs, unclear business value and inadequate risk controls as the causes. Anushree Verma, a senior director analyst at the firm, described most current agentic projects as early-stage experiments driven by hype and often misapplied.

The same release estimates that of the thousands of vendors claiming agentic capability, only around 130 offer something genuine. That figure belongs in every procurement conversation. A deployment built on a repackaged workflow tool inherits the cost of an agent and the capability of a script.

Where agentic AI money actually sat in early 2025 Gartner poll of 3,412 webinar attendees, January 2025 3,412 respondents Conservative investment 42% Waiting, or unsure 31% Significant investment 19% No investment 8% A webinar audience self-selects toward interest, so read this as an upper bound on adoption.
Four fifths of this audience had either invested conservatively or not committed at all. The cancellation forecast applies to a much smaller base of serious spending than the headline suggests.

Abandonment rose sharply, and that is a healthy signal

S&P Global Market Intelligence surveyed more than 1,000 respondents across North America and Europe. It found the share of businesses scrapping most of their AI initiatives rising to 42% from 17% a year earlier. The average organisation scrapped 46% of its proofs of concept before production.

A rising abandonment rate is the market learning to measure. An organisation that cancels 46% of its experiments is running a portfolio. One that cancels almost nothing is either extraordinarily good at selection or not checking at all.

The pilot-stage version of this question is handled separately in the breakdown of six agent pilot failure modes. That post covers which failure modes stop projects before production, and why the widely quoted 88% figure has no source. This post picks up where it stops, at the deployments that did ship.

Forrester reported in June 2026 that three quarters of enterprise leaders say they are adopting agentic AI. Only a small minority run it in meaningful production beyond chatbot-style deployments. The same analysis found more than half of enterprises reporting agent sprawl, even after adopting a recognised governance framework.

Sprawl is a cost word. It means agents running without a named owner, without a budget line and without anyone computing what they return. That is the precise condition under which a negative number goes unnoticed for a year.

Where this argument is weakest

Three things about the case above deserve to be said out loud. The first concerns a number in the brief for this post.

The 22% figure could not be verified

A figure circulating widely holds that 22% of agent deployments report negative return at twelve months. The claimed root-cause split is 41% unclear success criteria, 33% insufficient tool or data access and 26% evaluation drift. It is usually attributed to Forrester.

I could not trace it. Every restatement found during this research sat on an aggregator page. None linked to a Forrester report, press release or briefing with a stated sample. Forrester's own published material on agentic AI in 2026 contains no such figure. So this post does not use it, and neither should you until somebody produces the source.

Every survey here has a survivorship problem

Each of these datasets asks people inside organisations about outcomes those same people are partly responsible for. Self-reported return is a claim, not an audit. The direction of the bias is not obvious.

Respondents may overstate returns to protect a decision they championed. They may equally understate them, because unmeasured benefits are invisible to the person filling in a form. No published dataset found during this research resolves that.

The case that negative return is the correct price

The strongest argument against this framing is simple. A twelve-month accounting window is the wrong instrument for a capability build. Learning what an agent cannot do is worth paying for, and that lesson arrives as a cost.

I accept that case for a bounded number of experiments with a stated learning goal and a stated stop date. I do not accept it for a production deployment running for a year. At that point the organisation is not learning, it is subsidising. IBM's own respondents expect the subsidy to end: 85% told the firm they expect positive returns from scaled efficiency investments by 2027.

A twelve-month ledger you can actually run

The remedy for both forecast errors is the same short document. Write the ledger before approval. Run it again at month twelve against exactly the same lines.

Lines that decide the sign of the result
LineUsually in the business case?Where the real number lives
Licence or platform feeYesThe contract, including any minimum commitment
Inference at production volumeRarely, and usually at pilot ratesProvider usage console, measured over a full month at real volume
Human review and rework timeNoAsk the receiving team to log it for two weeks
Integration and data access engineeringPartly, and usually as a one-offEngineering time sheets across the first two quarters
Incident response and rollbackNoOperations ticket volume tagged to the deployment
Cost of removing it laterNoEstimate the workflow rebuild before you depend on the workflow
Benefit, against a recorded baselineAsserted, not measuredThe metric value from the month before launch, written down

The last row decides everything else. If it is empty at approval, the deployment cannot produce a positive result later. There will be nothing to compare against. The related economics are set out in the analysis of the AI unit economics margin trap.

One more discipline costs nothing. Write down the number below which you will stop, and the date you will check it, in the document that approves the spend. A stop condition agreed in advance is the only version of this decision that does not become political later.

Where measurable return has shown up, and what those cases share, is examined in the piece on who is actually making money from generative AI.

Frequently asked questions

What is negative ROI in AI?

Negative ROI in AI means the total cost of running a deployment over a period exceeded the value it produced in the same period. It differs from zero return, where a pilot produced nothing measurable and then stopped. A negative-return deployment keeps operating, so it keeps generating inference charges, supervision time and integration work while the benefit stays unproven.

What percentage of AI projects have negative ROI?

No published survey measures that directly, and figures circulating with a specific share should be checked against a primary source before use. What is measured is adjacent. IBM found 25% of AI initiatives delivered their expected return across 2,000 CEOs surveyed in 2025. S&P Global found 42% of businesses scrapping most AI initiatives, up from 17% the year before.

Why do AI agent deployments lose money?

Two forecast errors compound. The benefit is set before the buyer understands the value, which IBM found 64% of CEOs admit to. The cost is estimated at pilot scale and then runs at production scale, adding inference volume, human review time and integration work. Because both errors come from the same optimism, they add together rather than cancelling out.

How do you measure ROI on an AI agent?

Record the baseline value of the target metric in the month before launch, with the period it covers. Track seven cost lines rather than one: licence, inference at real volume, human review time, integration engineering, incident response, removal cost and any minimum contract commitment. Compare at twelve months, not at the quarter after launch, because most cost growth appears between months four and nine.

When should you cancel an AI project?

Cancel when the twelve-month ledger is negative and no line has a credible path to changing. Set that stop number and its review date in the approval document, before anyone is invested in the outcome. Gartner expects over 40% of agentic projects to be cancelled by the end of 2027, so treat cancellation as a normal portfolio outcome rather than a failure event.

Is AI cost overrun the main reason projects fail?

Cost is one of three causes Gartner names, alongside unclear business value and inadequate risk controls. In practice the benefit side does more damage, because an unmeasured benefit defaults to zero when finance reviews it. Forrester found fewer than one third of decision-makers able to tie AI value to financial growth, which makes almost any cost figure look unjustified by comparison.

Where to start this week

Pick your single largest AI deployment and answer one question about it: what was the number before you turned it on. Not the estimate, the recorded value of the metric it was meant to move, in the month before launch.

If that number exists, run the seven-line ledger above against it this quarter. If it does not exist, you have found the finding. Record the baseline now, and treat the last twelve months as unmeasurable rather than as a success.

Then put one date in the calendar: twelve months from launch, or three months from today if launch was long ago. Write the stop number next to it.

Related analysis

This piece covers deployments that shipped. For the stage before that, read the six failure modes that stop agent pilots reaching production. For the cost line that grows fastest, read why cheaper tokens produce larger bills.

References

  1. IBM Institute for Business Value, CEOs double down on AI while navigating enterprise hurdles, 6 May 2025. 2,000 CEOs, 33 countries, 24 industries, fieldwork February to April 2025. Used for the 25%, 16%, 64% and 85% figures.
  2. Forrester, 2026 Technology and Security Predictions, 28 October 2025. Used for the deferred AI spend prediction and the share of decision-makers who can tie AI value to financial growth.
  3. Forrester, The state of agentic AI in 2026, 3 June 2026. Used for adoption against production figures and the agent sprawl finding.
  4. Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation forecast, the causes, the January 2025 poll of 3,412 attendees and the vendor count.
  5. CIO Dive, AI project failure rates are on the rise, 14 March 2025, reporting S&P Global Market Intelligence Voice of the Enterprise research, more than 1,000 respondents in North America and Europe. Used for the 42% and 46% figures.
  6. BetterUp Labs and Stanford Social Media Lab, Workslop research, survey of 1,150 full-time United States desk workers, September 2025, first published in Harvard Business Review. Used for the $186 monthly figure and the two-hour resolution time.
  7. FinOps Foundation and Linux Foundation, State of FinOps 2026 survey, 19 February 2026. 1,192 respondents representing more than $83 billion in annual cloud spend. Used for the 98% and 31% figures.
  8. Cloudera, Data Readiness Index, 14 April 2026, fieldwork by Researchscape, 1,270 IT leaders, 22 January to 3 March 2026. Vendor-commissioned. Used for the data access and data confidence figures, both labelled as such in the text.

The weakest thing about this source base: every figure here is self-reported by people with a stake in the outcome, and none of it is an audited financial result. Two of the eight sources are forecasts rather than measurements, and one is vendor-commissioned. The waterfall figure is illustrative and contains no measured amounts.

SK
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading