From Sanskriti Khandelwal | Product & Market Analysis

Scaling AI Agents: What the 25% Who Reached Production Did Differently

On this page

Scaling AI agents past the pilot is rarer than vendor decks claim and more common than the failure headlines claim. Deloitte found 25% of organisations have moved 40% or more of their AI pilots into production. McKinsey found 23% scaling an agentic system in at least one function. The minority is roughly a quarter, and it shares five habits.

Key takeaways

  • The scaled minority is about a quarter, not a rounding error. Deloitte puts 25% of organisations past the 40% pilot-to-production mark. McKinsey puts 23% scaling an agentic system in at least one business function. BCG classes 5% as future-built and another 35% as scalers.
  • Process redesign is the one practice every dataset agrees on. McKinsey's high performers redesign workflows at 55%, against roughly 20% elsewhere. Only 21% of organisations using generative AI have redesigned any workflow at all.
  • Ownership is a measurable variable, not a platitude. Among the 26% of executives reporting accelerating year-on-year returns, 48% say decision authority for AI work is extremely clear. That is the cheapest item on the list to copy.
  • The scaled minority is still narrow, and still ungoverned. No more than 10% of respondents in any single business function say they are scaling agents, and only 21% report a mature governance model for them.
25%Organisations that moved 40% or more of their AI pilots into production. Source: Deloitte, January 2026.
23%Organisations scaling an agentic system in at least one function. Source: McKinsey, November 2025.
26%Executives whose AI returns are accelerating year on year. Source: Google Cloud, July 2026.

What "scaled" actually means in these surveys

Every headline share of scaled agent programmes measures something different. Fix that before comparing them, because the spread runs from 5% to 35% and most of the gap is definitional.

One survey counts organisations with any agent in production. Another counts those that moved a defined share of their pilot portfolio. A third counts those whose returns are accelerating rather than merely positive.

Three definitions, three numbers

Deloitte's State of AI in the Enterprise, published 21 January 2026, surveyed 3,235 director-to-C-suite leaders across 24 countries. It found 25% had moved 40% or more of their AI pilots into production. A further 54% expect to reach that threshold within three to six months, which is an intention rather than a result.

McKinsey's State of AI, published November 2025 from 1,993 respondents, found 23% actively scaling an agentic system in at least one business function. Another 39% were still experimenting. Note the bar: one function is not the enterprise, and no more than 10% of respondents in any given function report scaling agents there.

BCG surveyed 1,250 senior executives and AI decision makers across nine industries, assessed on 41 capabilities. It classed 5% as future-built, 35% as scalers and 60% as laggards. A third of the future-built group were using agents, against 12% of scalers and almost none of the laggards.

The 14% figure, and where it comes from

A widely circulated claim says only 14% of enterprises have scaled an agent to company-wide use, sourced to a March 2026 survey of 650 technology leaders. That survey is the publisher's own, run by a marketing agency. No methodology document, questionnaire or dataset has been released.

The same publisher is the origin point for the 88% agent pilot failure rate, which our analysis of agent pilot failure modes could not trace to any survey with a stated sample. I am not saying the 14% is wrong. I am saying no reader can check it, and a smaller number with a published fielding window is worth more than a sharper one without.

Five research programmes, five definitions of the scaled minority
SourceSampleWhat it measuredMinority share
MIT NANDA, Aug 202552 interviews, 153 surveys, 300 deploymentsPilots with measurable profit and loss impact5%
BCG, Sep 20251,250 executives, 9 industriesFirms rated future-built on 41 capabilities5% (plus 35% scalers)
McKinsey, Nov 20251,993 respondents, ~105 countriesScaling an agentic system in one or more functions23%
Deloitte, Jan 20263,235 leaders, 24 countriesMoved 40% or more of AI pilots into production25%
Google Cloud, Jul 20262,403 executives, run with NRGFinancial returns accelerating year on year26%

Shares are not comparable across rows. They sit together to show how much of the disagreement is a difference in the question asked.

How big is the minority? It depends entirely on the question Each bar answers a different question. None measures company-wide agent scale directly. MIT NANDA5%Pilots with P&L impact BCG future-built5%Rated on 41 capabilities McKinsey23%Scaling in one function Deloitte25%40%+ of pilots shipped Google Cloud26%Returns accelerating BCG scalers35%Value starting Samples range from 205 to 3,235 respondents.
Read the right-hand labels before the bars. The 5% and the 26% are not disagreeing with each other, because they are counting different things.

1. They redesigned the process, not just the tool

One practice appears in every dataset that separates the scaled minority from everyone else. It is not model choice, budget size or vendor. It is whether the organisation changed the process the agent runs inside.

McKinsey's high performers redesign workflows when deploying AI at a rate of 55%, against roughly 20% of everyone else. Across all organisations using generative AI, only 21% have redesigned any workflow at all. Nearly 80% are layering AI on top of processes that were designed for people.

I read that 21% as the most useful number in this entire literature. It says most AI programmes are software purchases wearing the costume of a transformation, and it explains the pilot graveyard better than any claim about model quality.

Deloitte lands in the same place from another angle. It found 30% redesigning key processes around AI, against 37% using AI at surface level with minimal process change. BCG puts it more bluntly in its January 2026 note on scaling AI with new processes rather than new tools.

The 10/20/70 split

BCG's guiding principle for resource allocation is 10/20/70. Ten percent of effort goes to algorithms, 20% to technology and data, and the remaining 70% to people and processes. Most programmes invert that ratio, and the inversion shows up in their results a year later.

This is why agent orchestration tooling rarely rescues a stalled programme. The tooling sits in the 20%. The reason the pilot stalled usually sits in the 70%.

What redesign looks like in one process

BCG published a worked case in January 2026. An industrial goods company rebuilt its quote-to-order process around agents rather than bolting agents onto the existing one. The target design has agents resolve around 70% of requests for quotation with no human involvement, 20% with some human intervention, and 10% routed to people as the most complex transactions.

The company expects to cut labour costs in that process by 30% to 40%. The plan runs across two releases over 15 to 18 months. That timeline is the part vendors leave out of the deck, and it is the part that decides whether your programme survives a budget cycle.

2. They put one named owner on it

Google Cloud's ROI of AI 2026 study, run with National Research Group across 2,403 executives, isolated a cohort it calls AI ROI Leaders. These are the 26% reporting returns that accelerate year on year, rather than merely increase. The distinction earns its keep, because 84% report increasing returns and that number separates almost nobody.

Among that cohort, 48% say ownership and decision-making authority for AI and agentic work is extremely clear. Nearly 50% say AI is embedded in core business processes and revenue streams, or is enabling new revenue. And 38% have capability development built into roles with required training.

Clear ownership is the least interesting item on that list and the one I would fix first. It costs a meeting, not a budget, and it is the only one of the three you can finish this quarter.

What the 26% with accelerating returns say about themselves Share of AI ROI Leaders reporting each practice. Google Cloud with National Research Group, 2,403 executives. 48% Ownership and authority described as extremely clear ~50% AI embedded in core processes and revenue 38% Required training built into job roles
None of these dials reaches half. Even inside the winning cohort, the practices that correlate with returns are minority behaviour.

Ownership means the budget line, not the org chart

Accountability without control is the common failure pattern. A named owner needs three specific things: the budget, the authority to stop the project, and an on-call obligation once the agent runs against live data. An owner with only the first is a sponsor, and sponsors do not fix incidents at 2am.

Google Cloud sells the products this study measures, so discount it accordingly. The methodology is published and the sample is large, which is more than most vendor research offers. Read the direction of the finding, not the decimal place.

3. They bought more than they built

The MIT NANDA research on generative AI in business drew on 52 executive interviews, 153 leader surveys and 300 public deployments. It found 95% of pilots produced no measurable profit and loss impact. The surviving 5% shared a pattern that had nothing to do with model selection.

They bought from vendors rather than building internally. They chose back-office friction over customer-facing showpieces. And they measured workflow change rather than licence adoption, which is the metric most programmes report because it is the one that always goes up.

Buying is faster. It is not cheaper. BCG puts vendor platforms at up to $1.5 million per use case or function, around three times the typical annual run cost of an in-house platform. What the money buys is time and somebody else's maintenance burden.

The integration question decides it

The build-or-buy answer usually turns on how deep the agent has to reach into your systems. A bought agent that cannot see your data is a demo. Standardised connection layers such as the Model Context Protocol have shifted this calculation, because integration is now less often the reason to build.

Our note on build against buy for coding agents works through the same trade in a narrower domain, and the conclusion holds here. Build where the process is the product. Buy where it is overhead.

4. They started where a wrong answer is cheap

Deloitte's own read of its data is that the companies seeing most success start with lower-risk use cases, build governance capability alongside, and scale deliberately. Gartner reaches the same conclusion from the opposite direction. It names escalating cost, unclear business value and inadequate risk controls as the reasons more than 40% of agentic AI projects will be cancelled by the end of 2027.

MIT's surviving 5% picked back-office friction for a reason worth stating plainly. A wrong invoice classification costs an hour of somebody's time. A wrong answer to a customer costs the account. Starting where errors are cheap buys you the iteration count you need, and iteration count is what turns a demo into a deployment.

That is also the honest argument against the showcase pilot. Customer-facing agents demo beautifully to a board and generate the incident profile that gets a programme paused.

Efficiency is the entry, not the destination

BCG found that among its leading firms, focusing on efficiency is the least common path, followed by only 10% of them. Efficiency is where they start, because it produces quick wins and proves the case for further investment. It is not where they stay.

That sequencing changes how you write the business case. An efficiency pilot that hits its number and then stops is a success by its own terms and a failure by the programme's. Write the second phase into the first proposal, or the win becomes the ceiling. Our analysis of who is actually making money on generative AI covers where those second phases have landed.

5. They funded fluency, not just the model

Deloitte recorded workforce access to sanctioned AI tools rising from under 40% to around 60% in a single year. Access is not adoption. A licence issued is not a workflow changed, and the gap between those two numbers is where most programme value quietly disappears.

The Google Cloud cohort treats this as a funded obligation rather than an enablement afterthought. Among ROI Leaders, 38% report comprehensive, ongoing capability development embedded into roles with required training. Required is the operative word. Optional training reaches the people who needed it least.

If your agent programme has a training budget smaller than its inference budget, that ratio is telling you which failure you are about to have.

Which practice is supported by which research programme Filled means measured. Outline means implied, not measured. Blank means not covered. Deloitte McKinsey BCG Google Cl. MIT 1. Redesign the process 2. One named owner 3. Buy more than build 4. Start where errors are cheap 5. Fund fluency Only practice 1 is measured by all five. Practices 2 and 3 rest on two sources each.
The blank cells are the useful part. A practice supported by one vendor study is a hypothesis, not a finding.

Governance arrived before the scale, not after

Deloitte found roughly 75% of organisations plan to deploy agentic AI within two years, and only 21% report a mature governance model for agents today. That is a gap of three and a half to one between intent and control.

Customisation plans make it worse. 85% of companies expect to customise agents to their own needs, and every customisation is a system nobody else can audit, patch or benchmark for you.

Gartner adds a supply-side warning that belongs in any vendor shortlist. It describes widespread agent washing, the rebranding of assistants, robotic process automation and chatbots as agents, and estimates only around 130 of the thousands of agentic vendors are real. The scaled minority did their diligence before procurement, not during the incident review.

The cost line most programmes fail to model is the per-task one. An agent that resolves a ticket for $0.40 and one that resolves it for $4.00 look identical in a demo and differ by an order of magnitude at volume. Our breakdown of the conversation maths behind agent pricing shows how quickly that compounds.

The five practices, the weak version, and the version that works
PracticeWeak versionVersion that works
1. Redesign the processAgent added to the existing workflowProcess rebuilt around what the agent can resolve unaided
2. One named ownerA steering committee and an executive sponsorOne person with the budget, a kill switch and an on-call rota
3. Buy more than buildInternal platform built before a use case existsBought where the process is overhead, built where it is the product
4. Start where errors are cheapCustomer-facing agent chosen because it demos wellBack-office task where a wrong answer costs an hour
5. Fund fluencyLicences issued, optional enablement offeredRequired training in the role description, with a budget line

Where this argument is weakest

Two problems sit under everything above, and the confident tone of the practice list should not hide them.

Every number here is self-reported

Deloitte, McKinsey, BCG and Google Cloud all asked executives what their organisations do. Nobody observed a deployment. Self-assessment against a term like "fundamentally redesigned the workflow" is exactly the kind of question people answer flatteringly, and the high performers have the strongest reason to answer it that way.

There is a sharper version of this objection. The practice list may describe what successful firms say about themselves after the fact, rather than what they did before the outcome was known.

Survivorship runs through all of it

These studies segment by outcome and then look backwards for shared traits. That method finds real patterns and it also finds coincidences, because the firms with accelerating AI returns are disproportionately the firms that were already well run. Process redesign capability is not something an AI programme creates.

The practice list is still worth copying. Hold it as a description of conditions that help rather than a causal recipe, and expect the copy to underperform the original.

Frequently asked questions

What percentage of companies have scaled AI agents company-wide?

No survey measures company-wide agent scale cleanly. The nearest measured figures are Deloitte's 25% of organisations that moved 40% or more of their AI pilots into production, and McKinsey's 23% actively scaling an agentic system in at least one business function. Within any single function, McKinsey puts the share scaling agents at no more than 10%. Company-wide scale is rarer than either number.

What do companies that successfully scale AI agents do differently?

Five practices recur across Deloitte, McKinsey, BCG, Google Cloud and MIT research. They redesign the process rather than adding an agent to it. They name one owner with budget authority. They buy more than they build. They start where a wrong answer is cheap, usually in the back office. And they fund user training as a line item rather than an afterthought.

Is the 14% AI agent scaling figure accurate?

It cannot be checked. The 14% comes from a March 2026 survey of 650 enterprise technology leaders run and published by a marketing agency, with no methodology document, questionnaire or dataset released. The same publisher originated the 88% agent pilot failure rate, which also has no traceable source. The figure may be right. Nothing available lets a reader verify it.

Should you build or buy AI agents for production?

MIT NANDA found the 5% of pilots that produced measurable profit impact mostly bought from vendors rather than building internally. Buying is faster, not cheaper. BCG puts vendor platforms at up to $1.5 million per use case, around three times the typical annual run cost of an in-house platform. Build when the process is your differentiator. Buy when it is not.

How long does it take to scale an AI agent to production?

Longer than a pilot timeline suggests. The worked example BCG published in January 2026 plans a redesigned quote-to-order process across two releases over 15 to 18 months. Deloitte found 54% of organisations expect to move 40% or more of their pilots into production within three to six months, which is a stated intention rather than a measured result. Plan in quarters.

What is the biggest blocker to AI agent production deployment?

Process design, on the available evidence. McKinsey found only 21% of organisations using generative AI have redesigned any workflow, so most are adding agents to processes built for people. Gartner names escalating cost, unclear business value and inadequate risk controls as the reasons more than 40% of agentic projects will be cancelled by the end of 2027. Model quality ranks below all four.

Where to start this week

Take the agent project furthest along in your organisation and draw its process twice. Draw it as it runs today. Then draw it as you would design it if the agent already worked and you were starting from a blank page.

If the two drawings look the same, you are in the 80% layering AI onto a process built for people, and no model upgrade will fix that. The gap between the two drawings is your actual scope, and it is almost always larger than the pilot charter admits.

Then write one name next to the second drawing. Not a committee, not a sponsor. The person who holds the budget, can stop the work, and gets the call when the agent does something expensive at 2am. If nobody will accept that name, you have learned more this week than any pilot result would have taught you.

If you take one thing

Ask whoever quotes you a scaling statistic which question the survey actually asked. The spread between 5% and 35% is almost entirely definitional, and the definition is where the argument is.

References

  1. Deloitte, From ambition to activation, 21 January 2026. State of AI in the Enterprise, 3,235 leaders, 24 countries, fielded August to September 2025. Used for the 25%, 54%, 30%, 37%, 21% and 85% figures and the tool access shift.
  2. McKinsey, The state of AI in 2025: agents, innovation, and transformation, November 2025. 1,993 respondents, roughly 105 countries. Used for the 23%, the 10% per-function ceiling and the 21% workflow redesign figure. The report PDF did not load during this research, so the 55% against 20% high-performer split is taken from published summaries and should be upgraded to a page reference.
  3. Forbes, Roughly 10% of enterprise functions use AI agents, McKinsey finds, 22 March 2026. Secondary reporting used to confirm the 23%, 39% and per-function figures.
  4. Google Cloud, How AI ROI Leaders prioritize investments, 28 July 2026. ROI of AI 2026 with National Research Group, 2,403 executives. Used for the 26%, 48%, near-50% and 38% figures. Vendor-commissioned research on products the vendor sells.
  5. BCG, AI leaders outpace laggards, 30 September 2025. 1,250 executives, nine industries, 41 capabilities. Used for the 5%, 35%, 60% and agent usage figures.
  6. BCG, Scaling AI requires new processes, not just new tools, 20 January 2026. Used for the 10/20/70 principle, the quote-to-order case, the 30% to 40% cost reduction, the 15 to 18 month timeline and the $1.5 million platform comparison. Single case, no sample.
  7. Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation forecast, the named causes and the agent washing estimate.
  8. Forbes, MIT finds 95% of GenAI pilots fail because companies avoid friction, 26 August 2025. Reporting MIT NANDA: 52 interviews, 153 surveys, 300 deployments. Used for the 95% and the surviving 5% pattern.

The weakest thing about this source base: five of the eight sources segment respondents by outcome, then look backwards for shared traits. That method cannot separate a cause from a symptom, and one of the five is published by a company selling what it measured.

SK
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Works in People and Culture at Wayground (Quizizz), and writes here on what AI actually does to how software teams work, hire and are measured.

Related reading