From Sanskriti Khandelwal | Product & Market Analysis

Support Teams After 60% Deflection: What Is Actually Left, and How to Staff It

On this page

Support deflection at 60% does not remove 60% of the work. It removes the fastest, most repetitive contacts and leaves a queue in which almost every case is an exception. Gartner now expects half the companies that cut service headcount for AI to rehire by 2027. The staffing model failed, not the technology.

Key takeaways

  • Deflection removes contacts and workload in very different proportions. If the automated 60% were also the fastest 60%, total handle minutes fall by roughly a third, and headcount follows the minutes rather than the contact count.
  • Only 20% of service leaders have actually cut agent headcount because of AI. Gartner surveyed 321 customer service and support leaders in October 2025 and found the layoff narrative running well ahead of the practice.
  • The tier-one queue was the training system, and deflection dismantled it. A peer-reviewed study of 5,179 support agents found AI raised novice output by 34% and experienced output barely at all, so the tool that clears junior work also removes the reason to hire juniors.
  • Rising average handle time after deflection is a pass, not a fail. Every metric calibrated on the old case mix now reads a harder queue as a performance problem, which is how competent teams get restructured for the wrong reason.
50%Share of companies that blamed headcount cuts on AI which Gartner expects to rehire for similar work by 2027. Source: Gartner, February 2026.
20%Service leaders who have actually reduced agent staffing because of AI, in a survey of 321. Source: Gartner, October 2025 fieldwork.
34%Productivity gain for novice agents given an AI assistant, against minimal gain for experienced agents. Source: Quarterly Journal of Economics, 2025.

What deflection actually removes from the queue

Deflection is the share of inbound contacts that never reach a person. Most teams report it as one percentage against total volume. That single number hides the thing that decides your staffing.

Contacts are not interchangeable units. A password reset and a disputed chargeback each count as one contact. They do not cost the same to serve, and they do not automate at the same rate.

Automation gets pointed at the intents that are easiest to specify, because those are the ones you can write a policy and a retrieval path for. Those are also the intents that were already cheapest to serve. Deflection is therefore never a random sample of your queue. It is drawn deliberately from the shallow end.

Deflection and resolution are different measures

A conversation that ends without a human is deflected. A conversation that ends with the customer's problem solved is resolved. The two sets overlap heavily and they are not the same set.

A customer who gives up and returns tomorrow on a different channel counts as a deflection today and a fresh contact tomorrow. Vendors have an incentive to report the first number. Buyers need the second, and the commercial consequences of that gap are worked through in the analysis of per-resolution pricing and what a vendor counts as a resolution.

Ask for the repeat-contact rate on deflected conversations over the following 7 days, split by intent. If nobody can produce it, your deflection rate is not a measurement. It is an assumption with a decimal point on it.

The easy contacts were also the fast contacts

Handle time and automatability are correlated because both are driven by the same underlying property: how much of the case can be specified in advance. A refund inside policy is short and automatable for the same reason. A billing dispute involving a partial refund, a broken integration and an angry finance director is neither.

So the average handle time of your human queue rises mechanically, without any individual agent getting slower.

The arithmetic: 60% of contacts is not 60% of the work

Take a queue of 10,000 contacts a month at an average handle time of 8 minutes. That is 80,000 minutes of human work.

Now deflect 60% of the contacts. Those 6,000 conversations averaged 5 minutes each, because they were the simple ones. You removed 30,000 minutes.

You did not remove 48,000 minutes. The remaining 4,000 contacts still carry 50,000 minutes between them, which is 12.5 minutes each. Your workload fell by 37.5%, not by 60%.

Deflecting 60% of contacts removes 37.5% of the minutes Illustrative arithmetic on a 10,000 contact queue. Not measured data. 80,000 min 10,000 contacts 30,000 min out 6,000 at 5 min 50,000 min 4,000 at 12.5 min Before deflection Automated away What is left Cut the team by 60% here. You then have 40% of the staff serving 62.5% of the workload.
Notice the third bar is not shorter than half the first. The contact count fell by 60% and the minutes by 37.5%.
The same queue, counted two ways
MeasureBefore deflectionAfter 60% deflectionChange
Contacts reaching a human10,0004,000Down 60%
Average handle time8.0 minutes12.5 minutesUp 56%
Total human handle minutes80,00050,000Down 37.5%
Defensible headcount cutBaselineAbout 37%Not 60%

These are illustrative figures, not measurements from any published dataset. The direction holds whenever automated intents are faster than the average.

The gap between 60 and 37.5 is where support restructuring goes wrong. A plan built on the contact count cuts 60% of the team and hands the survivors 62.5% of the original workload. It falls out of the arithmetic and shows up as attrition two quarters later.

Two inputs decide your own version of this number. The first is the pre-automation handle time of exactly the intents you automated, which your helpdesk already stores. The second is the observed handle time of what remains, measured after the rollout rather than forecast before it. I disagree with the common framing that deflection is a cost programme. It is a case-mix change that happens to have a cost effect, and treating it as the former is what produces the wrong headcount number.

Four kinds of work that do not deflect

The residual queue is not a smaller version of the old one. Each part of it demands something automation was never going to supply.

What survives automation, and why
Residual categoryWhy automation does not clear itWhat it needs from a person
Ambiguous intentThe customer cannot name the problem, so retrieval has nothing to match againstDiagnosis, not lookup
Exceptions across systemsThe fix crosses systems with no clean write path or no permission to writeAccess, and authority to act
High consequenceBeing wrong costs money, a regulator's attention or the accountJudgement, and someone accountable
Repair of the automationThe agent already failed on this contact and the customer arrives annoyedRecovery, plus a defect report

Work that cannot be specified in advance

The first three rows are familiar to anyone who has run an escalation desk. The fourth is new, and in most organisations it is unstaffed.

Every automated queue generates a stream of contacts that arrive already damaged. The customer has spent 6 minutes with an agent that did not understand them, and they now want both the original problem solved and an acknowledgement that the last 6 minutes were wasted. That is two jobs, and the second one did not exist in your old contact mix.

It also carries information nobody is collecting. The agent handling that contact is the only person in the company who can see exactly where the automation broke. Route those cases to a named owner with a weekly review, or the same failure runs for another quarter. Designing that path deliberately rather than discovering it is the point of treating the human loop as architecture rather than a fallback.

My position is that repair volume should be a reported line on the support dashboard from day one of any deployment. If you cannot separate it from ordinary escalations, you cannot tell an improving automation from a degrading one.

Your staffing model is calibrated on a mix that no longer exists

Support capacity planning rests on a small number of stable assumptions. Deflection breaks most of them at once, quietly, and none of them announce it.

Headcount by volume is now the wrong formula

Forecast contacts, divide by contacts per agent per hour, add shrinkage, staff to a service level. That model assumes contacts per hour is a property of the agent. After deflection it is mostly a property of the queue.

Rebuild the forecast in minutes rather than contacts. It is the same data and a different denominator, and it is the single highest-value change on this list.

Tiering breaks next. A three-rung structure of tier one, tier two and escalation assumes a broad base of simple work supporting a narrow top. Automate the base and you are left with a two-rung ladder that still has three sets of job descriptions, pay bands and promotion paths attached to it. That misalignment is the same one showing up a layer above, in what happens to middle management when agents absorb the coordination work.

Occupancy targets are the quiet danger. An 85% occupancy target on 5-minute lookups is demanding. The same target on 12-minute cases that each require diagnosis and judgement is a different job at the same number. I would drop the occupancy target by 5 to 10 points on the day a large deflection programme goes live, and revisit it once the residual handle time has settled.

The metrics invert, and the dashboard calls it failure

Support reporting is built on ratios that were meaningful against a stable case mix. Change the mix and each of them keeps producing a number while quietly changing what the number means.

The same metric, before and after a deflection programme
MetricWhat it measured beforeWhat it measures nowWhat to do
Average handle timeAgent efficiencyCase difficultyRemove it from individual targets, keep it as a mix diagnostic
CSATService qualityMostly case mixCompare within intent, never in aggregate across the change
Deflection rateAutomation coverageLittle on its ownAlways pair with the 7-day repeat contact rate
OccupancyCapacity utilisationBurnout exposureLower the target, then re-derive it from observed data

Handle time going up is a pass, not a fail

If your residual handle time did not rise after deflection, one of two things is true. Either the automation is taking hard cases it should not be taking, or it is closing conversations that were never resolved. Both are worth investigating, and neither is good news.

Treat a flat line as the anomaly and go looking for the reason.

Aggregate CSAT deserves the same scepticism. The contacts that left your human queue were the ones customers rated most highly, because they were quick and they worked. Removing them lowers the human average without any agent doing worse. Compare like with like or you will performance-manage people for a change you made yourself, which is a specific case of the wider problem in rewriting performance criteria once AI sits inside the work.

The apprenticeship problem nobody budgeted for

The post-deflection queue needs experienced agents, because everything left in it requires judgement. The uncomfortable part is that deflection removes the mechanism that produced experienced agents.

Tier one was never only a cost centre. It was where new hires learned the product, the systems, the tone and the edge cases, on contacts where being slow or wrong was cheap. Automate it and you have kept the requirement and deleted the training ground.

The assistant lifts beginners. It barely moves experts. Change in issues resolved per hour, 5,179 customer support agents. Brynjolfsson, Li and Raymond, QJE 2025. Novice and low skilled +34% All agents, average +14% Experienced and skilled Minimal effect reported by the authors The tool spreads the practices of able workers. It moves new agents down the experience curve.
Read this as a statement about who the tool replaces. It substitutes for the first year of learning, which is exactly the year you stopped paying for.

Brynjolfsson, Li and Raymond studied the staggered rollout of a conversational assistant across 5,179 customer support agents and found a 14% average gain in issues resolved per hour. The average conceals the finding that matters: 34% for novice and low-skilled workers, and minimal effect for experienced and highly skilled ones.

The assistant is a substitute for inexperience. It is not a substitute for expertise, and expertise is what the residual queue consumes.

The entry-level cohort is already thinner

This is not a forecast. It is visible in payroll data now. Stanford's Digital Economy Lab, using ADP administrative records through June 2026, reports that workers aged 22 to 25 in the most AI-exposed occupations show a 19% employment decline relative to less-exposed peers. The same paper finds no evidence of economy-wide displacement, and finds the decline is driven by reduced hiring rather than higher departures.

Customer service representatives sit squarely in the exposed group. The Bureau of Labor Statistics counts 2,666,000 of them in 2025 and projects a 5% decline through 2035. It still expects about 289,500 openings a year, all of them from replacement rather than growth. The occupation is shrinking and churning at the same time.

That combination is the trap. You will keep needing to fill seats, and the pool of people who learned the job on easy tickets will keep getting smaller. The identical dynamic is playing out one department over, in the breaking pipeline for junior developers, and support leaders have the advantage of watching it happen there first.

If I were running this, I would fund a deliberate training pipeline out of the deflection savings, in the same budget line, in the same quarter. Route a fixed share of automatable contacts to new hires on purpose.

What the public record shows about cutting too far

Three organisations have now run this experiment in public, and the sequence is consistent enough to plan against.

Two years of cutting support headcount in public Company announcements and analyst forecasts, February 2024 to February 2026 Feb 2024 Klarna assistant takes 2.3m chats in month one, said to equal 700 agents May 2025 Klarna reverses course and recruits humans, citing lower quality Sep 2025 Salesforce support goes from 9,000 to about 5,000 people Feb 2026 Gartner forecasts half of AI-linked cuts are rehired by 2027 Red markers are corrections. The first arrived 15 months after the first announcement. Sources: Klarna via Bloomberg and CX Dive, Salesforce Q2 FY2026 earnings call via CNBC, Gartner press release.
Klarna needed 15 months to discover the quality cost, which is longer than most deflection programmes run before the headcount decision is made.

Klarna deployed an OpenAI-based assistant in February 2024. It handled 2.3 million conversations in the first month, work the company described as equivalent to 700 full-time agents, and it resolved cases in about 2 minutes against 11 for humans. In May 2025 chief executive Sebastian Siemiatkowski told Bloomberg the company had gone too far. The published quote is that investing in the quality of human support is the way of the future for the company. Klarna began recruiting people again, while the assistant still handled about two thirds of inquiries.

Salesforce is the more instructive case, because it did not reverse. Marc Benioff confirmed on the Q2 FY2026 earnings call that customer support headcount had gone from about 9,000 to about 5,000, with AI agents handling 1.5 million inquiries over 9 months. He also said plainly that a great deal cannot be resolved by the agent and has to be escalated to humans, and the company reported an even split between AI-handled and human-handled interactions. The vendor comparison behind those deployments is broken down in the review of Agentforce, Fin and Breeze as support agents.

Gartner has moved in the same direction across four releases. In March 2025 it forecast that agentic AI would autonomously resolve 80% of common customer service issues by 2029. By September 2025 it was predicting that no Fortune 500 company would have fully eliminated human customer service by 2028. In February 2026 it added that half of the companies attributing cuts to AI will rehire for similar functions under different job titles by 2027.

I think that rehiring is mostly a repricing rather than a reversal. The roles come back at a higher skill level, with different titles, and often at higher pay, because the work that remains is genuinely harder. That is a cost outcome very different from the original business case. The function-by-function version of that sum is in the analysis of where agent payback actually lands.

Where this argument is weakest

Three things here would not survive a determined challenge.

No public dataset measures the residual

The core claim, that residual handle time rises after deflection, is a mechanism argument rather than a measured one. I could not find a single credible public dataset that reports pre-deflection and post-deflection handle time for the same queue, split by intent.

The arithmetic in this post is therefore illustrative, and labelled as such in the figure and the table. The premise, that automated intents are faster than average, is almost certainly true in most queues, and you should confirm it in yours rather than take it from here.

Salesforce's own numbers cut the other way

Salesforce reported that its remaining human agents handled roughly the same number of conversations at nearly identical satisfaction scores after the cut. If the residual queue were meaningfully harder per case, you would expect one of those two numbers to move.

Several explanations fit. The routing may be sending more than tier-one work to the agent. The team may have absorbed the difficulty without it showing up in either metric yet. But the honest reading is that a large, well-instrumented deployment has published figures that do not support the concentration story, and one counterexample from a company with real data outweighs a lot of reasoning.

The third weakness is scale. Gartner's own survey found only 20% of service leaders have actually cut agent headcount because of AI. A post about what happens at 60% deflection is addressing a minority of organisations. Most teams are approaching that situation rather than living in it.

Frequently asked questions

What is a good support deflection rate in 2026?

There is no credible public benchmark. The percentages circulating on this question come from vendor blogs and aggregator posts whose figures disagree by wide margins, and they define deflection inconsistently. Measure true resolution instead, meaning the share of conversations that end without a human and without the customer returning within 7 days. That number is comparable across your own quarters, which is the only comparison you need.

Does AI deflection reduce support headcount proportionally?

No, and assuming it does is the most common planning error. Automation clears the intents that are easiest to specify, and those are typically the shortest. Deflecting 60% of contacts that averaged 5 minutes against a queue average of 8 removes about 37% of the handle minutes, not 60%. Build capacity plans on total handle minutes rather than contact counts, and the gap disappears from your forecast.

Why does average handle time go up after AI deflection?

Because the case mix changed, not because agents got slower. The contacts that automation absorbs are short, well-defined and repetitive. What remains is ambiguous, cross-system, high-consequence or already broken by a failed automated attempt. A rise in residual handle time is the expected outcome. If handle time stays flat, check whether the automation is closing conversations that were never actually resolved.

How should you restructure a support team after AI deflection?

Move capacity planning from contacts to minutes, lower the occupancy target while the residual mix settles, and collapse tier one and tier two into a single diagnostic role with real system authority. Then staff a named owner for contacts that arrive after an automated failure, because that stream is new and carries the only reliable signal about where the automation is breaking.

Are companies rehiring customer service staff they cut for AI?

Gartner predicts that by 2027, half the companies that attributed headcount reduction to AI will rehire staff for similar functions under different job titles. Klarna publicly reversed in May 2025 after quality complaints and resumed recruiting human agents. The pattern is less a reversal than a repricing, since roles tend to return at a higher skill level and a higher cost than the ones removed.

What happens to entry-level support jobs when AI handles tier one?

They shrink, and with them the training path to senior roles. Stanford's Digital Economy Lab finds workers aged 22 to 25 in AI-exposed occupations down 19% relative to less-exposed peers, driven by reduced hiring rather than departures. The Bureau of Labor Statistics projects a 5% decline in customer service representative employment through 2035 while still expecting about 289,500 replacement openings a year.

Where to start this week

One query and one decision, in that order.

The query: pull handle time by intent for the 90 days before your automation went live, then split those intents into the ones the agent now handles and the ones it does not. Multiply each group by its volume. You now have the real minutes removed and the real minutes remaining, and your capacity plan has a denominator that reflects the work. If you have not deployed yet, run the same query anyway.

The decision: pick the share of automatable contacts you will deliberately route to new hires, and write it into next quarter's budget as a training cost rather than an efficiency loss. The number can be small. What matters is that somebody has decided it on purpose, because the alternative is discovering in 18 months that you have a queue of hard cases and nobody who learned the job.

Related on this site

The vendor side of this question is covered in the comparison of Agentforce, Fin and Breeze. The escalation design that decides how much reaches a person is in the piece on human-in-the-loop as architecture.

References

  1. Gartner, Half of companies that cut customer service staff due to AI will rehire by 2027, 3 February 2026. Used for the 50% rehiring forecast and the October 2025 survey of 321 service leaders.
  2. Gartner via Businesswire, None of the Fortune 500 will have fully eliminated human customer service by 2028, 10 September 2025. Used for the Fortune 500 prediction and Kathy Ross's comments.
  3. Erik Brynjolfsson, Danielle Li and Lindsey Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025. Used for the 5,179 agent sample and the 14%, 34% and minimal-effect findings.
  4. Erik Brynjolfsson, Bharat Chandar and Ruyu Chen, Canaries in the Coal Mine, Stanford Digital Economy Lab, revised 12 August 2026. Used for the 19% relative decline among workers aged 22 to 25.
  5. US Bureau of Labor Statistics, Customer Service Representatives, Occupational Outlook Handbook. Used for 2025 employment, the 2025 to 2035 projection and annual openings.
  6. CNBC, Salesforce CEO confirms 4,000 layoffs because I need less heads with AI, 2 September 2025. Used for the 9,000 to 5,000 support headcount figure.
  7. CX Dive, Salesforce still sees a place for live customer service agents after massive cuts, 4 September 2025. Used for the escalation comment and the 1.5 million inquiries figure.
  8. CX Dive, Klarna changes its AI tune and again recruits humans for customer service, 9 May 2025. Used for all Klarna figures and the Siemiatkowski quote, sourced from a Bloomberg interview.

Weakest thing about this source base: the central claim that residual handle time rises after deflection has no measured public dataset behind it, and the arithmetic used to illustrate it is constructed rather than observed. Two Gartner figures come from press release text the newsroom would not serve to an automated fetch, taken as reported rather than from the page itself. The Klarna quotes are secondary coverage of a Bloomberg interview.

SK
Sanskriti Khandelwal
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading