From Sanskriti Khandelwal | Product & Market Analysis
Scaling AI Agents: What the 25% Who Reached Production Did Differently
On this page
Scaling AI agents past the pilot is rarer than vendor decks claim and more common than the failure headlines claim. Deloitte found 25% of organisations have moved 40% or more of their AI pilots into production. McKinsey found 23% scaling an agentic system in at least one function. The minority is roughly a quarter, and it shares five habits.
Key takeaways
- The scaled minority is about a quarter, not a rounding error. Deloitte puts 25% of organisations past the 40% pilot-to-production mark. McKinsey puts 23% scaling an agentic system in at least one business function. BCG classes 5% as future-built and another 35% as scalers.
- Process redesign is the one practice every dataset agrees on. McKinsey's high performers redesign workflows at 55%, against roughly 20% elsewhere. Only 21% of organisations using generative AI have redesigned any workflow at all.
- Ownership is a measurable variable, not a platitude. Among the 26% of executives reporting accelerating year-on-year returns, 48% say decision authority for AI work is extremely clear. That is the cheapest item on the list to copy.
- The scaled minority is still narrow, and still ungoverned. No more than 10% of respondents in any single business function say they are scaling agents, and only 21% report a mature governance model for them.
What "scaled" actually means in these surveys
Every headline share of scaled agent programmes measures something different. Fix that before comparing them, because the spread runs from 5% to 35% and most of the gap is definitional.
One survey counts organisations with any agent in production. Another counts those that moved a defined share of their pilot portfolio. A third counts those whose returns are accelerating rather than merely positive.
Three definitions, three numbers
Deloitte's State of AI in the Enterprise, published 21 January 2026, surveyed 3,235 director-to-C-suite leaders across 24 countries. It found 25% had moved 40% or more of their AI pilots into production. A further 54% expect to reach that threshold within three to six months, which is an intention rather than a result.
McKinsey's State of AI, published November 2025 from 1,993 respondents, found 23% actively scaling an agentic system in at least one business function. Another 39% were still experimenting. Note the bar: one function is not the enterprise, and no more than 10% of respondents in any given function report scaling agents there.
BCG surveyed 1,250 senior executives and AI decision makers across nine industries, assessed on 41 capabilities. It classed 5% as future-built, 35% as scalers and 60% as laggards. A third of the future-built group were using agents, against 12% of scalers and almost none of the laggards.
The 14% figure, and where it comes from
A widely circulated claim says only 14% of enterprises have scaled an agent to company-wide use, sourced to a March 2026 survey of 650 technology leaders. That survey is the publisher's own, run by a marketing agency. No methodology document, questionnaire or dataset has been released.
The same publisher is the origin point for the 88% agent pilot failure rate, which our analysis of agent pilot failure modes could not trace to any survey with a stated sample. I am not saying the 14% is wrong. I am saying no reader can check it, and a smaller number with a published fielding window is worth more than a sharper one without.
| Source | Sample | What it measured | Minority share |
|---|---|---|---|
| MIT NANDA, Aug 2025 | 52 interviews, 153 surveys, 300 deployments | Pilots with measurable profit and loss impact | 5% |
| BCG, Sep 2025 | 1,250 executives, 9 industries | Firms rated future-built on 41 capabilities | 5% (plus 35% scalers) |
| McKinsey, Nov 2025 | 1,993 respondents, ~105 countries | Scaling an agentic system in one or more functions | 23% |
| Deloitte, Jan 2026 | 3,235 leaders, 24 countries | Moved 40% or more of AI pilots into production | 25% |
| Google Cloud, Jul 2026 | 2,403 executives, run with NRG | Financial returns accelerating year on year | 26% |
Shares are not comparable across rows. They sit together to show how much of the disagreement is a difference in the question asked.
1. They redesigned the process, not just the tool
One practice appears in every dataset that separates the scaled minority from everyone else. It is not model choice, budget size or vendor. It is whether the organisation changed the process the agent runs inside.
McKinsey's high performers redesign workflows when deploying AI at a rate of 55%, against roughly 20% of everyone else. Across all organisations using generative AI, only 21% have redesigned any workflow at all. Nearly 80% are layering AI on top of processes that were designed for people.
I read that 21% as the most useful number in this entire literature. It says most AI programmes are software purchases wearing the costume of a transformation, and it explains the pilot graveyard better than any claim about model quality.
Deloitte lands in the same place from another angle. It found 30% redesigning key processes around AI, against 37% using AI at surface level with minimal process change. BCG puts it more bluntly in its January 2026 note on scaling AI with new processes rather than new tools.
The 10/20/70 split
BCG's guiding principle for resource allocation is 10/20/70. Ten percent of effort goes to algorithms, 20% to technology and data, and the remaining 70% to people and processes. Most programmes invert that ratio, and the inversion shows up in their results a year later.
This is why agent orchestration tooling rarely rescues a stalled programme. The tooling sits in the 20%. The reason the pilot stalled usually sits in the 70%.
What redesign looks like in one process
BCG published a worked case in January 2026. An industrial goods company rebuilt its quote-to-order process around agents rather than bolting agents onto the existing one. The target design has agents resolve around 70% of requests for quotation with no human involvement, 20% with some human intervention, and 10% routed to people as the most complex transactions.
The company expects to cut labour costs in that process by 30% to 40%. The plan runs across two releases over 15 to 18 months. That timeline is the part vendors leave out of the deck, and it is the part that decides whether your programme survives a budget cycle.
2. They put one named owner on it
Google Cloud's ROI of AI 2026 study, run with National Research Group across 2,403 executives, isolated a cohort it calls AI ROI Leaders. These are the 26% reporting returns that accelerate year on year, rather than merely increase. The distinction earns its keep, because 84% report increasing returns and that number separates almost nobody.
Among that cohort, 48% say ownership and decision-making authority for AI and agentic work is extremely clear. Nearly 50% say AI is embedded in core business processes and revenue streams, or is enabling new revenue. And 38% have capability development built into roles with required training.
Clear ownership is the least interesting item on that list and the one I would fix first. It costs a meeting, not a budget, and it is the only one of the three you can finish this quarter.
Ownership means the budget line, not the org chart
Accountability without control is the common failure pattern. A named owner needs three specific things: the budget, the authority to stop the project, and an on-call obligation once the agent runs against live data. An owner with only the first is a sponsor, and sponsors do not fix incidents at 2am.
Google Cloud sells the products this study measures, so discount it accordingly. The methodology is published and the sample is large, which is more than most vendor research offers. Read the direction of the finding, not the decimal place.
3. They bought more than they built
The MIT NANDA research on generative AI in business drew on 52 executive interviews, 153 leader surveys and 300 public deployments. It found 95% of pilots produced no measurable profit and loss impact. The surviving 5% shared a pattern that had nothing to do with model selection.
They bought from vendors rather than building internally. They chose back-office friction over customer-facing showpieces. And they measured workflow change rather than licence adoption, which is the metric most programmes report because it is the one that always goes up.
Buying is faster. It is not cheaper. BCG puts vendor platforms at up to $1.5 million per use case or function, around three times the typical annual run cost of an in-house platform. What the money buys is time and somebody else's maintenance burden.
The integration question decides it
The build-or-buy answer usually turns on how deep the agent has to reach into your systems. A bought agent that cannot see your data is a demo. Standardised connection layers such as the Model Context Protocol have shifted this calculation, because integration is now less often the reason to build.
Our note on build against buy for coding agents works through the same trade in a narrower domain, and the conclusion holds here. Build where the process is the product. Buy where it is overhead.
4. They started where a wrong answer is cheap
Deloitte's own read of its data is that the companies seeing most success start with lower-risk use cases, build governance capability alongside, and scale deliberately. Gartner reaches the same conclusion from the opposite direction. It names escalating cost, unclear business value and inadequate risk controls as the reasons more than 40% of agentic AI projects will be cancelled by the end of 2027.
MIT's surviving 5% picked back-office friction for a reason worth stating plainly. A wrong invoice classification costs an hour of somebody's time. A wrong answer to a customer costs the account. Starting where errors are cheap buys you the iteration count you need, and iteration count is what turns a demo into a deployment.
That is also the honest argument against the showcase pilot. Customer-facing agents demo beautifully to a board and generate the incident profile that gets a programme paused.
Efficiency is the entry, not the destination
BCG found that among its leading firms, focusing on efficiency is the least common path, followed by only 10% of them. Efficiency is where they start, because it produces quick wins and proves the case for further investment. It is not where they stay.
That sequencing changes how you write the business case. An efficiency pilot that hits its number and then stops is a success by its own terms and a failure by the programme's. Write the second phase into the first proposal, or the win becomes the ceiling. Our analysis of who is actually making money on generative AI covers where those second phases have landed.
5. They funded fluency, not just the model
Deloitte recorded workforce access to sanctioned AI tools rising from under 40% to around 60% in a single year. Access is not adoption. A licence issued is not a workflow changed, and the gap between those two numbers is where most programme value quietly disappears.
The Google Cloud cohort treats this as a funded obligation rather than an enablement afterthought. Among ROI Leaders, 38% report comprehensive, ongoing capability development embedded into roles with required training. Required is the operative word. Optional training reaches the people who needed it least.
If your agent programme has a training budget smaller than its inference budget, that ratio is telling you which failure you are about to have.
Governance arrived before the scale, not after
Deloitte found roughly 75% of organisations plan to deploy agentic AI within two years, and only 21% report a mature governance model for agents today. That is a gap of three and a half to one between intent and control.
Customisation plans make it worse. 85% of companies expect to customise agents to their own needs, and every customisation is a system nobody else can audit, patch or benchmark for you.
Gartner adds a supply-side warning that belongs in any vendor shortlist. It describes widespread agent washing, the rebranding of assistants, robotic process automation and chatbots as agents, and estimates only around 130 of the thousands of agentic vendors are real. The scaled minority did their diligence before procurement, not during the incident review.
The cost line most programmes fail to model is the per-task one. An agent that resolves a ticket for $0.40 and one that resolves it for $4.00 look identical in a demo and differ by an order of magnitude at volume. Our breakdown of the conversation maths behind agent pricing shows how quickly that compounds.
| Practice | Weak version | Version that works |
|---|---|---|
| 1. Redesign the process | Agent added to the existing workflow | Process rebuilt around what the agent can resolve unaided |
| 2. One named owner | A steering committee and an executive sponsor | One person with the budget, a kill switch and an on-call rota |
| 3. Buy more than build | Internal platform built before a use case exists | Bought where the process is overhead, built where it is the product |
| 4. Start where errors are cheap | Customer-facing agent chosen because it demos well | Back-office task where a wrong answer costs an hour |
| 5. Fund fluency | Licences issued, optional enablement offered | Required training in the role description, with a budget line |
Where this argument is weakest
Two problems sit under everything above, and the confident tone of the practice list should not hide them.
Every number here is self-reported
Deloitte, McKinsey, BCG and Google Cloud all asked executives what their organisations do. Nobody observed a deployment. Self-assessment against a term like "fundamentally redesigned the workflow" is exactly the kind of question people answer flatteringly, and the high performers have the strongest reason to answer it that way.
There is a sharper version of this objection. The practice list may describe what successful firms say about themselves after the fact, rather than what they did before the outcome was known.
Survivorship runs through all of it
These studies segment by outcome and then look backwards for shared traits. That method finds real patterns and it also finds coincidences, because the firms with accelerating AI returns are disproportionately the firms that were already well run. Process redesign capability is not something an AI programme creates.
The practice list is still worth copying. Hold it as a description of conditions that help rather than a causal recipe, and expect the copy to underperform the original.
Frequently asked questions
What percentage of companies have scaled AI agents company-wide?
No survey measures company-wide agent scale cleanly. The nearest measured figures are Deloitte's 25% of organisations that moved 40% or more of their AI pilots into production, and McKinsey's 23% actively scaling an agentic system in at least one business function. Within any single function, McKinsey puts the share scaling agents at no more than 10%. Company-wide scale is rarer than either number.
What do companies that successfully scale AI agents do differently?
Five practices recur across Deloitte, McKinsey, BCG, Google Cloud and MIT research. They redesign the process rather than adding an agent to it. They name one owner with budget authority. They buy more than they build. They start where a wrong answer is cheap, usually in the back office. And they fund user training as a line item rather than an afterthought.
Is the 14% AI agent scaling figure accurate?
It cannot be checked. The 14% comes from a March 2026 survey of 650 enterprise technology leaders run and published by a marketing agency, with no methodology document, questionnaire or dataset released. The same publisher originated the 88% agent pilot failure rate, which also has no traceable source. The figure may be right. Nothing available lets a reader verify it.
Should you build or buy AI agents for production?
MIT NANDA found the 5% of pilots that produced measurable profit impact mostly bought from vendors rather than building internally. Buying is faster, not cheaper. BCG puts vendor platforms at up to $1.5 million per use case, around three times the typical annual run cost of an in-house platform. Build when the process is your differentiator. Buy when it is not.
How long does it take to scale an AI agent to production?
Longer than a pilot timeline suggests. The worked example BCG published in January 2026 plans a redesigned quote-to-order process across two releases over 15 to 18 months. Deloitte found 54% of organisations expect to move 40% or more of their pilots into production within three to six months, which is a stated intention rather than a measured result. Plan in quarters.
What is the biggest blocker to AI agent production deployment?
Process design, on the available evidence. McKinsey found only 21% of organisations using generative AI have redesigned any workflow, so most are adding agents to processes built for people. Gartner names escalating cost, unclear business value and inadequate risk controls as the reasons more than 40% of agentic projects will be cancelled by the end of 2027. Model quality ranks below all four.
Where to start this week
Take the agent project furthest along in your organisation and draw its process twice. Draw it as it runs today. Then draw it as you would design it if the agent already worked and you were starting from a blank page.
If the two drawings look the same, you are in the 80% layering AI onto a process built for people, and no model upgrade will fix that. The gap between the two drawings is your actual scope, and it is almost always larger than the pilot charter admits.
Then write one name next to the second drawing. Not a committee, not a sponsor. The person who holds the budget, can stop the work, and gets the call when the agent does something expensive at 2am. If nobody will accept that name, you have learned more this week than any pilot result would have taught you.
If you take one thing
Ask whoever quotes you a scaling statistic which question the survey actually asked. The spread between 5% and 35% is almost entirely definitional, and the definition is where the argument is.
References
- Deloitte, From ambition to activation, 21 January 2026. State of AI in the Enterprise, 3,235 leaders, 24 countries, fielded August to September 2025. Used for the 25%, 54%, 30%, 37%, 21% and 85% figures and the tool access shift.
- McKinsey, The state of AI in 2025: agents, innovation, and transformation, November 2025. 1,993 respondents, roughly 105 countries. Used for the 23%, the 10% per-function ceiling and the 21% workflow redesign figure. The report PDF did not load during this research, so the 55% against 20% high-performer split is taken from published summaries and should be upgraded to a page reference.
- Forbes, Roughly 10% of enterprise functions use AI agents, McKinsey finds, 22 March 2026. Secondary reporting used to confirm the 23%, 39% and per-function figures.
- Google Cloud, How AI ROI Leaders prioritize investments, 28 July 2026. ROI of AI 2026 with National Research Group, 2,403 executives. Used for the 26%, 48%, near-50% and 38% figures. Vendor-commissioned research on products the vendor sells.
- BCG, AI leaders outpace laggards, 30 September 2025. 1,250 executives, nine industries, 41 capabilities. Used for the 5%, 35%, 60% and agent usage figures.
- BCG, Scaling AI requires new processes, not just new tools, 20 January 2026. Used for the 10/20/70 principle, the quote-to-order case, the 30% to 40% cost reduction, the 15 to 18 month timeline and the $1.5 million platform comparison. Single case, no sample.
- Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation forecast, the named causes and the agent washing estimate.
- Forbes, MIT finds 95% of GenAI pilots fail because companies avoid friction, 26 August 2025. Reporting MIT NANDA: 52 interviews, 153 surveys, 300 deployments. Used for the 95% and the surviving 5% pattern.
The weakest thing about this source base: five of the eight sources segment respondents by outcome, then look backwards for shared traits. That method cannot separate a cause from a symptom, and one of the five is published by a company selling what it measured.
Related reading