From Sanskriti Khandelwal | Product & Market Analysis
Sunk Cost and the AI Pilot: McKinsey's 37% Has Not Moved, Budgets Have
On this page
37% of respondents to McKinsey's 2026 State of AI survey say AI contributes any EBIT at all to their organisation. That is about the same share as a year earlier. Over the same period, budgets rose and overruns got funded. The sunk cost fallacy explains most of that gap, and an AI pilot is close to the ideal conditions for it.
Key takeaways
- Reported AI earnings impact is flat while spending is not. McKinsey's 2026 survey of 1,719 respondents found 37% attribute any EBIT impact to AI, about the same as 2025, and only 6% attribute 5% or more of EBIT.
- Overruns are mostly funded, not stopped. In Futurum's 2H 2026 survey of 1,636 technology decision makers, 46.9% were over their AI budget, and only 17.2% of those paused or reduced an initiative in response.
- Sunk cost is a forecasting error, not a moral one. Money already spent cannot be recovered by spending more, yet decades of decision research show people keep investing more in the courses of action they personally started.
- The cheapest fix is structural, not motivational. Re-approve every pilot as if it were new, price it only on the money still to be spent, and give the continuation decision to someone who did not sponsor the original.
The short answer
Stop an AI pilot when you would not approve it today as a new proposal. Price it only on the money still to be spent, and judge it on evidence of a path to EBIT rather than user enthusiasm. Spend already made is irrelevant to that decision. Have someone other than the original sponsor make the call.
McKinsey State of AI 2026: the 37% that did not move
McKinsey published "The state of AI in 2026: On the road to ROI" on 25 August 2026. The title is optimistic. The headline number underneath it is not.
According to The Register's account of the report, 37% of respondents attribute at least some EBIT impact to AI use, about the same share as in the 2025 survey. EBIT is earnings before interest and taxes, the operating profit line. "At least some" means any non-zero amount. The bar is set about as low as a survey question can set it, and nearly two thirds of respondents still do not clear it.
The top of the distribution is flat too. McKinsey's AI high performers are respondents who attribute at least 5% of EBIT to AI and describe the impact as significant. That group was 6% of the sample again in 2026.
What did move in the same survey
Almost everything upstream of profit rose. Among respondents at organisations above $1 billion in revenue, 40% said they are scaling AI agents, against 27% the year before. 80% of respondents who use AI in their roles said it improved their individual productivity.
The report's own summary line, quoted by The Register, is the most honest sentence in it. "Organizations' conviction in AI is growing faster than the immediate financial returns they can attribute to it." Conviction rising faster than returns is the precise behavioural signature of sunk cost. It is not proof, but it is the pattern you would expect if it were operating.
AI pilot ROI stalled, and AI budgets moved instead
Returns are flat. Spending is not, and three independent surveys published in 2026 point the same way.
Planned spend nearly doubled
KPMG's Q1 2026 AI Quarterly Pulse surveyed 237 US leaders at companies above $1 billion in revenue between 17 February and 17 March 2026. They projected average AI spending of $207 million over the next 12 months, which KPMG described as nearly double the figure from the same period a year earlier. In KPMG's industry modules, 96% of tech leaders and 94% of banking leaders said AI would stay a top investment priority even in a recession.
Secondary coverage of the McKinsey report by BIP Capital, citing the report's exhibits, puts the share of respondents who expect AI investment to rise next year at 60%. I could not open the McKinsey PDF to confirm that figure directly, so treat it as reported rather than verified.
Overruns get absorbed, not questioned
Futurum's 2H 2026 CIO and Technology Buyers survey was published 10 September 2026. It is the closest thing in the public record to a direct observation of sunk cost at budget level. Of 1,636 global enterprise technology decision makers, 46.9% said actual AI spend exceeded plan: 35.6% moderately and 11.3% substantially. Only 5.6% were below plan.
Then look at what the 767 over-plan organisations did about it. 47.6% asked for more budget. 43.3% absorbed the overrun and planned to settle it later. Only 17.2% paused or reduced the AI initiative. The categories overlap, so they do not sum to 100%, but the ranking is unambiguous.
None of this proves any individual pilot should have stopped. Some overruns are good news, because usage beat the forecast. What it shows is that the default response to new negative information is to add money, and that default is what the sunk cost literature predicts.
The sunk cost fallacy in AI projects, as a mechanism
A sunk cost is money, time or reputation already spent that cannot be recovered whatever you decide next. The fallacy is letting it influence the next decision. The correct comparison only ever looks forward: the remaining cost of continuing against the value continuing will produce.
Hal Arkes and Catherine Blumer set out the effect in "The psychology of sunk cost", published in Organizational Behavior and Human Decision Processes in 1985. Their work included a field study of theatre season ticket buyers, testing whether what people had paid changed how often they attended. A decade earlier, Barry Staw's 1976 paper "Knee-deep in the big muddy" ran a role-play experiment with business students. Those who were personally responsible for an initial investment that went badly committed more further resources to it.
Jeffrey Sleesman and colleagues published a meta-analytic review of the determinants of escalation of commitment in the Academy of Management Journal in 2012. Escalation of commitment is the organisational version of the same error: continuing a failing course of action because of what has already gone into it.
Three channels that keep a pilot alive
The first channel is self-justification. The sponsor who championed a pilot reads every ambiguous result as early evidence of success, because the alternative reading is that their judgement failed. Staw's experiment is about exactly this: responsibility, not information, drove the extra commitment.
The second is completion pull. A pilot that is "nearly there" feels cheaper to finish than to abandon. In practice, "nearly there" is often a description of the demo, not of the operating result. The remaining distance to production is usually the expensive part, a point explored in the tests that separate a proof of concept from a production system.
The third is social exposure. Once a pilot has been announced to a board or a team, stopping it is a public reversal. The cost of stopping is borne by a named person. The cost of continuing is spread across a budget line nobody reads closely.
Each channel alone produces a modest bias. Together, across a portfolio of 20 or 30 pilots, they produce a predictable shape: very few stops, many renewals, and a slow drift of spend toward initiatives whose main achievement is still existing. That shape matches the Futurum response pattern above more closely than any competing explanation I can find.
Why AI pilots are unusually exposed to sunk cost
AI pilots have three features that make escalation worse than in other technology programmes.
The evidence is real, and it is the wrong evidence
An AI pilot reliably generates positive user sentiment. McKinsey's 80% productivity figure is self-reported, and I have no reason to think it is false. The trouble is that "my work feels faster" does not convert into EBIT unless someone redesigns the workflow, changes the headcount plan or wins revenue that would otherwise have been lost.
That gives every sponsor a supply of true, favourable, irrelevant evidence. It is the perfect fuel for self-justification, because nobody has to lie. Futurum's February 2026 survey suggests buyers are starting to notice. Among 830 IT decision makers, the share naming productivity as their number one AI ROI metric fell from 23.8% to 18.0%. Direct financial impact rose to 21.7% of first-place answers.
Costs arrive as a meter, not a cheque
Traditional software is bought once and renewed annually. Most AI spend now arrives as consumption: tokens, credits, seats with usage caps. A meter never presents a single moment where someone signs a large number, so it never forces a re-decision. In the McKinsey survey, 20% of respondents said AI operating costs had constrained their use, which tells you the meter is large enough to feel. Why usage bills keep climbing even as unit prices fall is covered in the analysis of cheap tokens and rising AI bills.
The strategic story makes stopping look like retreat
Nobody is praised for cancelling an AI pilot in 2026. KPMG's recession figures show how deep the commitment runs: AI is positioned as the thing you keep funding when everything else is cut. That is a defensible strategic position for a portfolio. It is a terrible default for any single pilot inside it, because it turns "stop this one" into "retreat from AI". The two decisions need to be separated in every review.
What each 2026 AI survey can and cannot prove
The numbers above are only as good as the questions behind them.
| Survey | Sample and window | What it measures well | What it cannot tell you |
|---|---|---|---|
| McKinsey, State of AI 2026 | 1,719 respondents, global. Secondary coverage citing the methodology note gives fieldwork as 4 May to 8 June 2026 across 97 nations. | Trends in self-reported adoption and attribution, year on year | Actual EBIT. Attribution is the respondent's judgement, not an audited figure. |
| Futurum, 2H 2026 CIO and Technology Buyers | 1,636 global enterprise technology decision makers. Fielding dates not published. | Budget position and the response to overruns | Whether overruns were justified by usage. Futurum is a commercial analyst firm, and the release is a summary of a paid report. |
| KPMG, AI Quarterly Pulse Q1 and Q3 2026 | 237 (Q1) and 314 (Q3) US leaders at companies with revenue above $1 billion. Fielded 17 February to 17 March and 24 July to 25 August 2026. | Planned spend and leadership sentiment at large US companies | Smaller firms, non-US firms, or realised spend. Planned figures are intentions. |
KPMG's Q3 2026 Pulse found that 58% of leaders report measurable business value from AI, but only 37% name stronger financial performance as a value they have seen. That 37% matches McKinsey's 37% exactly. It is a coincidence of two different questions asked of two different samples, and I would not stack them as corroboration. What both say, independently, is that financial results trail every other kind of reported value. Where the measurable returns have actually landed is traced in the piece on who is making money from generative AI.
When to stop an AI project: the fresh-money rule
The six kill criteria for AI projects covered the thresholds to set at approval. This rule is the companion piece. It governs the moment a pilot comes back asking for its next tranche, which is where sunk cost does its damage.
The rule fits in one sentence. At every funding gate, ask whether you would approve this pilot today if it arrived as a new proposal, at its remaining cost, with the evidence it now has. If the answer is no, it stops or it shrinks. Money already spent is not admissible on either side of the argument.
| Gate element | Weak version | Version that resists sunk cost |
|---|---|---|
| Cost on the slide | Total invested to date, plus next tranche | Remaining cost to reach the next decision only. Spend to date removed from the deck. |
| Evidence accepted | User satisfaction, usage counts, demo quality | A measured change in a cost or revenue line, against a baseline recorded before launch |
| Who decides | The sponsor, presenting to peers | A reviewer who did not approve the original pilot |
| Default if undecided | Funding continues | Funding lapses at the gate date unless re-approved |
| Allowed outcomes | Continue or cancel | Scale, narrow to the one workflow that showed results, or stop |
Why the default has to flip
The most important row is the fourth. In the Futurum data, the default response to an overrun was more money. A gate where funding lapses unless someone actively re-approves it reverses the burden of proof. The pilot has to earn its next quarter.
I would set the gate at 90 days for most pilots. Long enough to record a baseline and see a change. Short enough that the sponsor has not yet stacked a year of public commitment on top of it.
Narrow is the outcome most teams forget
Most pilots that fail the fresh-money test do not fail everywhere. One team or one workflow usually shows a real result while the rest show enthusiasm. Narrowing the pilot to that slice is not a compromise. It is often the only version that would pass a fresh approval. Spreading a pilot wide before it works in one place is one of the patterns described in the common failure modes of agent pilots.
Change who decides, not just what they decide on
The single strongest piece of field evidence on de-escalation is not from technology. Barry Staw, Sigal Barsade and Kenneth Koput tracked 132 California banks over 9 years and published the results in the Journal of Applied Psychology in 1997. Turnover among senior bank executives predicted both provisions for loan losses and write-offs of bad loans. The reverse did not hold: provisions and write-offs did not predict executive turnover.
In plain terms, problem loans got recognised after the people who had made them left. They simply had no personal stake in the original decision being right.
You do not need to replace your executives to borrow that effect. You need to separate the person who approved a pilot from the person who renews it. A rotating review panel works. So does a finance partner with authority to let funding lapse. What does not work is the sponsor presenting to a committee of fellow sponsors, each of whom has a pilot of their own up for renewal next month.
If you change only one thing after reading this, change the reviewer. Thresholds can be argued with. A reviewer with nothing to defend is much harder to argue with.
Where this argument is weakest
Four objections are strong enough to state at full force.
Flat aggregate numbers can hide real progress. McKinsey's sample changes every year, and more organisations are now deep enough into AI to be asked about EBIT at all. A flat 37% could mean new entrants without returns are diluting incumbents who now have them. The survey design does not let me rule that out. Michael Chui, a coauthor, told The Register it "should not be surprising that it has taken time", pointing to patterns seen with earlier technologies.
Some continued spending is rational option value. A pilot can be worth funding past a weak quarter because it teaches an organisation something it will need later, such as how to evaluate models or how to change a workflow. Sunk cost reasoning and option reasoning look identical from outside. The difference is whether the team can name what it is learning and when that learning will be worth something.
Self-reported EBIT attribution is a weak measure in both directions. Respondents may overstate impact to justify spend. They may also understate it, because AI gains often show up as avoided hiring that no ledger records. I am treating 37% as a ceiling-ish signal of what people can see, not as the true figure.
The survey evidence does not observe individual pilots. Futurum's 17.2% describes organisations reacting to budget overruns, not decisions about specific projects with known outcomes. Linking that pattern to sunk cost is an inference from decades of experimental and field research, not a finding in the 2026 data. A competing explanation, that buyers rationally expect costs to fall and returns to rise, fits some of the same numbers.
None of these objections rescues the habit of funding a pilot because of what it has already cost. They do argue for calling the result of the fresh-money gate a judgement, not a calculation.
Frequently asked questions
What is the sunk cost fallacy in AI projects?
It is continuing to fund an AI project because of the money, time or reputation already spent on it, rather than because of the value still to come. Past spend cannot be recovered whatever you decide. The only valid comparison is the remaining cost of continuing against the remaining benefit. AI pilots are especially exposed because they produce favourable user sentiment even when no profit impact appears.
What did the McKinsey State of AI 2026 report find about EBIT impact?
McKinsey's 2026 survey of 1,719 respondents, published 25 August 2026, found that 37% attribute at least some EBIT impact to AI use, about the same share as in 2025. The AI high performer group, respondents attributing 5% or more of EBIT to AI with significant impact, stayed at 6%. Meanwhile 80% of AI users reported higher personal productivity, which shows how little of that reaches operating profit.
When should you stop an AI project?
Stop it when you would not approve it today as a new proposal, priced only on its remaining cost and judged on measured change in a cost or revenue line. Ignore spend to date entirely. If one workflow shows results and the rest do not, narrow the pilot rather than cancel it. Have someone who did not sponsor the original make the decision, and let funding lapse by default.
How do you measure AI pilot ROI properly?
Record a baseline for one cost or revenue line before the pilot starts, then measure the change in that line over a fixed period. User satisfaction and usage counts are inputs, not returns. Futurum's February 2026 survey of 830 IT decision makers found productivity falling as the top ROI metric, from 23.8% to 18.0%, while direct financial impact rose to 21.7% of first-place answers.
Why do companies keep funding AI pilots that show no results?
Three forces combine. Sponsors read ambiguous results favourably because stopping implies their judgement failed. Usage-based billing never forces a single large re-approval. And stopping an AI pilot can look like retreating from AI itself. Futurum's 2026 data shows the pattern: of organisations over their AI budget, 47.6% asked for more money and only 17.2% paused or reduced an initiative.
Does replacing the decision maker reduce sunk cost bias?
Field evidence suggests it does. Staw, Barsade and Koput studied 132 California banks over 9 years and found that senior executive turnover predicted later provisions and write-offs on problem loans, while write-offs did not predict turnover. New decision makers without a stake in the original choice recognised losses sooner. Inside a company, the practical version is separating the pilot sponsor from the renewal reviewer.
How to run your first fresh-money review
Pick the three AI pilots with the largest renewal coming up this quarter. For each one, write a single page that contains only the remaining cost to the next decision and the measured change in one cost or revenue line. Delete every reference to money already spent before anyone sees it.
Then hand those three pages to someone who approved none of the three pilots, and give that person the authority to let funding lapse. Whatever they decide, you will learn how much of your portfolio was running on its own history.
Related in this series
Set the thresholds before you start with the six kill criteria for AI projects, then check whether a reported win is real with the forensic test for AI ROI case studies.
References
- The Register, McKinsey says enterprise AI is finally 'on the road to ROI', 25 August 2026. Used for the 37% EBIT, 6% high performer, 80% productivity, 20% cost constraint, agent scaling, sample size and the Chui quotes.
- BIP Capital, The six percent: reading McKinsey's State of AI in 2026 as an investor, 15 September 2026. Secondary source, used only for fieldwork dates, country count and the 60% investment expectation, all citing the McKinsey report's own pages.
- Futurum Research, 46.9% of Enterprises Report AI Spend Over Budget in 2H 2026, 10 September 2026. Used for budget position and responses to overruns.
- Futurum Research, Enterprise AI ROI Shifts as Agentic Priorities Surge, 17 February 2026. Used for the change in top ROI metric.
- KPMG US, AI Quarterly Pulse Survey, Q1 2026, 31 March 2026. Used for planned spend and recession priority figures.
- KPMG US, AI Quarterly Pulse Survey, Q3 2026, 24 September 2026. Used for measurable value and financial performance shares.
- Staw, B. M., Barsade, S. G. and Koput, K. W., Escalation at the credit window, Journal of Applied Psychology, 82(1), 130 to 142, 1997. Used for the bank turnover finding.
- Staw, B. M., Knee-deep in the big muddy, Organizational Behavior and Human Performance, 16(1), 27 to 44, 1976. Used for the personal responsibility finding.
The weakest part of this source base is that McKinsey's own report could not be opened during research, so every McKinsey figure here comes through Tier 2 press coverage or a secondary investor note. All survey figures are self-reported and current as of 8 October 2026.
Related reading