From Shubhi K | Product & Market Analysis

GenAI ROI: Why 95% of Pilots Show No Return, and Where the Money Actually Lands

On this page

A widely cited MIT report found that about 95% of enterprise generative AI pilots produced no measurable profit and loss impact. The number is real and the study is preliminary. The more useful finding underneath it is that most organisations never recorded a baseline, so GenAI ROI could not have been measured even if it existed.

Key takeaways

  • The 95% figure comes from a preliminary, non peer-reviewed study. MIT's Project NANDA based it on 300 plus public deployments, 52 executive interviews and about 150 survey responses, gathered in the first half of 2025.
  • Most failures are measurement failures. Roughly half of GenAI budget went to sales and marketing while clearer returns appeared in back-office automation, per the same report.
  • The money today is concentrated upstream. Chip, cloud and model suppliers are capturing the spend, while most buyers cannot yet show a P&L line that moved.
  • The gap is between pilot and production, not between demo and disappointment. The average organisation scrapped 46% of its AI proofs of concept before they reached production, and 42% abandoned most initiatives outright, up from 17% a year earlier.
95%Share of enterprise GenAI pilots showing no measurable P&L impact. Source: MIT Project NANDA, The GenAI Divide, July 2025. Preliminary, not peer reviewed.
42%Share of companies that abandoned most AI initiatives, up from 17% a year earlier. Source: S&P Global Market Intelligence, 2025.
46%Average share of AI proofs of concept scrapped before reaching production. Source: S&P Global Market Intelligence, 2025.

What the 95% number actually says, and what it does not

The statistic comes from The GenAI Divide, published by MIT's Project NANDA in July 2025. It reported that roughly 95% of integrated enterprise AI pilots showed no measurable effect on profit and loss, while about 5% were extracting significant value.

The sample, and the fair criticism of it

The finding rests on analysis of more than 300 publicly disclosed AI deployments, 52 structured executive interviews and roughly 150 survey responses. The data was gathered in the first half of 2025.

That is a real research base and a small one. It is not a census of enterprise AI, and the authors described the work as preliminary.

The criticism is fair. Several critiques have landed. The study was not peer reviewed. The success definition was narrow, requiring measurable return at roughly six months. The interview base of 52 is thin for a claim applied to every industry.

Treat 95% as directional rather than precise. The direction has held up across other work, including a separate finding that 42% of companies abandoned most AI initiatives in 2025, up from 17% a year earlier.

FindingSource and yearWhat limits it
About 95% of GenAI pilots show no measurable P&L impactMIT Project NANDA, 2025Preliminary, not peer reviewed, 52 interviews, 6 month window
42% of companies abandoned most AI initiatives, up from 17%S&P Global Market Intelligence, 2025Self-reported, and abandonment is not the same as failure
At least 30% of generative AI projects abandoned after proof of concept by end of 2025Gartner, July 2024 forecastA forecast, and the measured outcome exceeded it
Average of 46% of AI proofs of concept scrapped before productionS&P Global Market Intelligence, 2025Self-reported, and abandonment reasons vary widely

Three separate studies pointing the same way is stronger evidence than any one of them. None of them establishes the precise figure.

Four studies, four different questions, one direction Read the label under each bar. They are not measuring the same thing. 95% no measurableP&L returnMIT NANDA, 202552 interviews, preliminary 46% of proofs of conceptscrappedS&P Global, 20251,006 enterprises 42% abandoned mostinitiativesS&P Global, 2025was 17% in 2024 30% predictedabandonmentGartner, July 2024forecast, not measured
The tallest bar is also the weakest evidence. The two middle bars have the largest sample and get quoted least.

Why most pilots return nothing

The money went where the excitement was

The MIT report found roughly half of GenAI budgets going to sales and marketing, while the clearer returns appeared in unglamorous back-office automation. Attention and payoff were pointed in different directions.

The pattern repeats across categories. Budget goes to the function that will demo the result at a board meeting, while the countable saving sits in reconciliation, data cleanup and internal coordination that nobody presents.

And nobody wrote down the starting number. This is the failure I see most often. A team deploys a tool, output feels faster, and six months later nobody can prove anything because the before state was never recorded.

You cannot compute a return without a baseline. If your business case says "improve efficiency" without a number and a date, the pilot has already failed its own test.

Generic tools stall at the second week

The report described tools that do not retain feedback or adapt to how an operation actually runs. They demo well and get abandoned once the workflow needs context.

This is why a general assistant often shows no measurable return while a narrow, embedded tool does. The difference is not model quality. It is whether the tool knows your data and your process.

So who is actually making money

Follow the cash and it clusters in three places, only one of which is the buyer.

The suppliers

Chip, power and cloud suppliers are capturing the spending directly. Four hyperscalers alone guided to roughly $725 billion of 2026 capital expenditure, most of it flowing to a narrow set of vendors. That build and its funding are examined in the piece on 2026 AI capex.

The model labs

Revenue at the two largest labs has grown at rates with no software precedent, reaching tens of billions annualised by mid 2026. Their costs have grown alongside, and neither has demonstrated sustained operating profitability.

Some of that revenue also originates from investors who are simultaneously suppliers, which makes it harder to read as pure customer demand. That structure is unpacked in the breakdown of the AI circular deals.

The narrow deployments

The third group is quiet and rarely written about. These are single-workflow deployments with a recorded baseline, a named owner and a metric that finance already tracks. They do not produce press releases. They produce a smaller invoice.

The common shape is narrow. One workflow, one metric finance already tracks, one person accountable for the number, and a deliberate decision not to expand until the first result holds for two quarters. That sequence is slower than a company-wide rollout and it is the only one that reliably produces a figure you can defend.

What a measured return actually looks like

The 5% is not a mystery. Where returns have been documented, three characteristics repeat, and none of them is about model quality.

It is one workflow, not a rollout

The documented wins are narrow. One process, one metric, one team. The MIT authors found the clearest returns in back-office automation rather than in the sales and marketing functions that absorbed roughly half the budget.

Scope choice appears to matter more than vendor choice. That is an uncomfortable finding for anyone whose first step was a vendor selection process.

The metric already existed before the tool

Successful deployments measure something finance was already tracking. Cost per ticket, days to close, error rate, hours billed. Nobody invents a new metric to prove the tool worked, because a metric created after the fact convinces no one.

This is also why back-office automation outperformed sales and marketing. Back-office processes have countable outputs. Marketing impact does not, at least not inside a six-month measurement window.

The failure happens between pilot and production

S&P Global surveyed 1,006 enterprises across North America and Europe. The average organisation scrapped 46% of its AI proofs of concept before production. Cost, data privacy and security were the most cited obstacles.

Gartner had predicted the direction. In July 2024 it forecast that at least 30% of generative AI projects would be abandoned after proof of concept by the end of 2025. It named poor data quality, weak risk controls, escalating costs and unclear business value. The measured outcome was worse.

Where all of this evidence is weak

A post criticising other people's statistics has to say what is wrong with the ones it relies on. Here it is.

Every figure above is self-reported. Executives surveyed about their own AI programmes have obvious incentives, in both directions, and none of these studies observed a controlled comparison. There is no version of any of these companies running the same year without the software.

Abandonment figures are also self-defined. A company deciding what counts as "most of our AI initiatives" is making a judgement call, and the threshold is not standardised across respondents.

Treat the whole evidence base as directional. The consistent direction across independent studies is worth more than any individual number, and no individual number here deserves to be quoted as precise.

The baseline problem, and how to fix it this week

Most AI ROI questions are unanswerable because of a decision made before the tool arrived. No baseline was recorded.

Fixing it does not require a project. It requires one hour and a spreadsheet with four columns. The metric, the value today, the date, and what you are deliberately excluding.

Do this before the next tool goes in. If a tool is already live and you have no baseline, take the earliest reliable period you can reconstruct and label it as reconstructed. An imperfect, honest baseline beats none.

Expect the reconstruction to be uncomfortable. Most teams discover the metric they want was never recorded consistently, which is itself the finding worth taking to the next vendor conversation.

A measurement template you can copy

ColumnWhat goes in itWorked example
MetricSomething finance already tracksCost per resolved support ticket
Baseline valueThe number before deployment$14.20 per ticket
PeriodStart and end date of the baseline1 April to 30 June, before deployment
ScopeWhich teams, products or clients are includedOne team, tier 1 tickets only
ExclusionsWhat you are deliberately leaving outImplementation time, licence cost
OwnerOne named person accountable for the numberHead of support operations

The exclusions row is the one people skip and the one that decides whether the number survives scrutiny. If licence cost is not in your calculation, say so rather than letting a reader assume it was.

The six fields that make a number defensible Record all six before deployment. A number missing any one of them will not survive scrutiny. METRICCost per resolved ticket BASELINE VALUE$14.20 PERIOD1 Apr to 30 Jun, pre-deployment SCOPEOne team, tier 1 tickets only EXCLUSIONSLicence cost, implementation time OWNERHead of support operations The red field is the one people skip. If licence cost is not in your calculation, say so. Letting a reader assume it was included is how a real saving becomes an indefensible claim.
Values shown are an illustrative worked example. The point is the shape of the record, not the numbers in it.

Where buying beats building, and where it does not

The MIT report noted that internally built tools reached production far less often than purchased ones. That is a real finding and it is not universal.

Building wins when the workflow is genuinely unique to you and you have engineering capacity to maintain it for years. Buying wins when the workflow is common across your industry, because the vendor amortises the integration work across many customers.

Most workflows are more common than the teams running them believe. Ticket triage, document review, data reconciliation and reporting look broadly similar across companies in the same sector, which is why an internal build often costs more over three years than the licence it replaced. If your process is genuinely unusual, that logic reverses and you should build.

Frequently asked questions

Is it true that 95% of AI projects fail?

The 95% figure comes from MIT's Project NANDA report of July 2025, which found that share of enterprise generative AI pilots showed no measurable profit and loss impact. It is preliminary and was not peer reviewed, and the sample included 52 executive interviews and around 150 survey responses. Read it as directional. Separate work found 42% of companies abandoning most AI initiatives in 2025.

Why do most GenAI pilots show no ROI?

Three reasons appear repeatedly. Budget went to visible functions such as sales and marketing rather than to back-office workflows where returns were clearer. No baseline was recorded before deployment, so return could not be calculated afterwards. And generic tools that do not retain context stall once a real workflow demands it.

How do I measure AI ROI properly?

Record a baseline before deployment using a metric your finance team already tracks. Note the value, the exact period, the scope of what is included, and what you are deliberately excluding such as licence cost and implementation time. Assign one named owner. Without those five fields, any number you produce later will not survive scrutiny.

Who is actually making money from AI right now?

Spending is concentrated upstream. Chip, power and cloud suppliers capture most of it, with four hyperscalers guiding to roughly $725 billion of 2026 capital expenditure. The largest model labs have grown revenue to tens of billions annualised without demonstrating sustained operating profit. Among buyers, measured returns cluster in narrow, embedded workflows rather than in general assistants.

Should we build our own AI tools or buy them?

Build when the workflow is genuinely specific to your business and you can maintain the system for years. Buy when the workflow is common across your industry, because a vendor spreads the integration cost across many customers. MIT's report found internally built tools reached production far less often than purchased ones, though that finding does not hold in every context.

How many AI projects actually reach production?

Fewer than most budgets assume. S&P Global surveyed 1,006 enterprises in 2025. The average organisation scrapped 46% of its AI proofs of concept before production, and 42% of companies abandoned most of their AI initiatives that year, up from 17%. Gartner had forecast at least 30% abandonment after proof of concept, so the measured outcome ran ahead of the prediction.

Where to start this week

Two actions, both small.

First, pick the AI tool you spend the most on and try to write its baseline from memory. The metric, the number before you bought it, and the date. If you cannot, you have found your real problem and it is not the tool.

Second, run the audit nobody runs. Take the dataset your AI tool depends on and count how much of it is duplicated, unusable, or accessible only to the person who created it. That number usually reframes the next purchase decision entirely, because it explains results the model was never going to fix.

Part three of three

Part one covers the $725 billion 2026 build. Part two maps who pays whom in the circular deals.

References

  1. MIT Project NANDA, The GenAI Divide: State of AI in Business 2025, July 2025, as reported by Fortune. Used for the 95% figure and the budget allocation finding.
  2. Virtualization Review, MIT report finds most AI business investments fail, August 2025. Used for the study's methodology and sample.
  3. Forbes, MIT finds 95% of GenAI pilots fail because companies avoid friction, 26 August 2025. Used for the pilot-to-production analysis.
  4. CIO Dive, AI project failure rates are on the rise, March 2025, reporting S&P Global Market Intelligence's survey of 1,006 respondents. Used for the 42% and 46% figures.
  5. CNBC, Hyperscalers face higher capex scrutiny, July 2026. Used for the 2026 capital expenditure figure.
  6. Forbes, OpenAI and Anthropic are testing two very different AI business models, 21 May 2026. Used for lab revenue run rates.

The MIT figure is preliminary and was not peer reviewed. It is reported here with its stated limitations rather than as a settled measurement.

SK
Shubhi K
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading