From Sanskriti Khandelwal | Product & Market Analysis
Support Teams After 60% Deflection: What Is Actually Left, and How to Staff It
On this page
Support deflection at 60% does not remove 60% of the work. It removes the fastest, most repetitive contacts and leaves a queue in which almost every case is an exception. Gartner now expects half the companies that cut service headcount for AI to rehire by 2027. The staffing model failed, not the technology.
Key takeaways
- Deflection removes contacts and workload in very different proportions. If the automated 60% were also the fastest 60%, total handle minutes fall by roughly a third, and headcount follows the minutes rather than the contact count.
- Only 20% of service leaders have actually cut agent headcount because of AI. Gartner surveyed 321 customer service and support leaders in October 2025 and found the layoff narrative running well ahead of the practice.
- The tier-one queue was the training system, and deflection dismantled it. A peer-reviewed study of 5,179 support agents found AI raised novice output by 34% and experienced output barely at all, so the tool that clears junior work also removes the reason to hire juniors.
- Rising average handle time after deflection is a pass, not a fail. Every metric calibrated on the old case mix now reads a harder queue as a performance problem, which is how competent teams get restructured for the wrong reason.
What deflection actually removes from the queue
Deflection is the share of inbound contacts that never reach a person. Most teams report it as one percentage against total volume. That single number hides the thing that decides your staffing.
Contacts are not interchangeable units. A password reset and a disputed chargeback each count as one contact. They do not cost the same to serve, and they do not automate at the same rate.
Automation gets pointed at the intents that are easiest to specify, because those are the ones you can write a policy and a retrieval path for. Those are also the intents that were already cheapest to serve. Deflection is therefore never a random sample of your queue. It is drawn deliberately from the shallow end.
Deflection and resolution are different measures
A conversation that ends without a human is deflected. A conversation that ends with the customer's problem solved is resolved. The two sets overlap heavily and they are not the same set.
A customer who gives up and returns tomorrow on a different channel counts as a deflection today and a fresh contact tomorrow. Vendors have an incentive to report the first number. Buyers need the second, and the commercial consequences of that gap are worked through in the analysis of per-resolution pricing and what a vendor counts as a resolution.
Ask for the repeat-contact rate on deflected conversations over the following 7 days, split by intent. If nobody can produce it, your deflection rate is not a measurement. It is an assumption with a decimal point on it.
The easy contacts were also the fast contacts
Handle time and automatability are correlated because both are driven by the same underlying property: how much of the case can be specified in advance. A refund inside policy is short and automatable for the same reason. A billing dispute involving a partial refund, a broken integration and an angry finance director is neither.
So the average handle time of your human queue rises mechanically, without any individual agent getting slower.
The arithmetic: 60% of contacts is not 60% of the work
Take a queue of 10,000 contacts a month at an average handle time of 8 minutes. That is 80,000 minutes of human work.
Now deflect 60% of the contacts. Those 6,000 conversations averaged 5 minutes each, because they were the simple ones. You removed 30,000 minutes.
You did not remove 48,000 minutes. The remaining 4,000 contacts still carry 50,000 minutes between them, which is 12.5 minutes each. Your workload fell by 37.5%, not by 60%.
| Measure | Before deflection | After 60% deflection | Change |
|---|---|---|---|
| Contacts reaching a human | 10,000 | 4,000 | Down 60% |
| Average handle time | 8.0 minutes | 12.5 minutes | Up 56% |
| Total human handle minutes | 80,000 | 50,000 | Down 37.5% |
| Defensible headcount cut | Baseline | About 37% | Not 60% |
These are illustrative figures, not measurements from any published dataset. The direction holds whenever automated intents are faster than the average.
The gap between 60 and 37.5 is where support restructuring goes wrong. A plan built on the contact count cuts 60% of the team and hands the survivors 62.5% of the original workload. It falls out of the arithmetic and shows up as attrition two quarters later.
Two inputs decide your own version of this number. The first is the pre-automation handle time of exactly the intents you automated, which your helpdesk already stores. The second is the observed handle time of what remains, measured after the rollout rather than forecast before it. I disagree with the common framing that deflection is a cost programme. It is a case-mix change that happens to have a cost effect, and treating it as the former is what produces the wrong headcount number.
Four kinds of work that do not deflect
The residual queue is not a smaller version of the old one. Each part of it demands something automation was never going to supply.
| Residual category | Why automation does not clear it | What it needs from a person |
|---|---|---|
| Ambiguous intent | The customer cannot name the problem, so retrieval has nothing to match against | Diagnosis, not lookup |
| Exceptions across systems | The fix crosses systems with no clean write path or no permission to write | Access, and authority to act |
| High consequence | Being wrong costs money, a regulator's attention or the account | Judgement, and someone accountable |
| Repair of the automation | The agent already failed on this contact and the customer arrives annoyed | Recovery, plus a defect report |
Work that cannot be specified in advance
The first three rows are familiar to anyone who has run an escalation desk. The fourth is new, and in most organisations it is unstaffed.
Every automated queue generates a stream of contacts that arrive already damaged. The customer has spent 6 minutes with an agent that did not understand them, and they now want both the original problem solved and an acknowledgement that the last 6 minutes were wasted. That is two jobs, and the second one did not exist in your old contact mix.
It also carries information nobody is collecting. The agent handling that contact is the only person in the company who can see exactly where the automation broke. Route those cases to a named owner with a weekly review, or the same failure runs for another quarter. Designing that path deliberately rather than discovering it is the point of treating the human loop as architecture rather than a fallback.
My position is that repair volume should be a reported line on the support dashboard from day one of any deployment. If you cannot separate it from ordinary escalations, you cannot tell an improving automation from a degrading one.
Your staffing model is calibrated on a mix that no longer exists
Support capacity planning rests on a small number of stable assumptions. Deflection breaks most of them at once, quietly, and none of them announce it.
Headcount by volume is now the wrong formula
Forecast contacts, divide by contacts per agent per hour, add shrinkage, staff to a service level. That model assumes contacts per hour is a property of the agent. After deflection it is mostly a property of the queue.
Rebuild the forecast in minutes rather than contacts. It is the same data and a different denominator, and it is the single highest-value change on this list.
Tiering breaks next. A three-rung structure of tier one, tier two and escalation assumes a broad base of simple work supporting a narrow top. Automate the base and you are left with a two-rung ladder that still has three sets of job descriptions, pay bands and promotion paths attached to it. That misalignment is the same one showing up a layer above, in what happens to middle management when agents absorb the coordination work.
Occupancy targets are the quiet danger. An 85% occupancy target on 5-minute lookups is demanding. The same target on 12-minute cases that each require diagnosis and judgement is a different job at the same number. I would drop the occupancy target by 5 to 10 points on the day a large deflection programme goes live, and revisit it once the residual handle time has settled.
The metrics invert, and the dashboard calls it failure
Support reporting is built on ratios that were meaningful against a stable case mix. Change the mix and each of them keeps producing a number while quietly changing what the number means.
| Metric | What it measured before | What it measures now | What to do |
|---|---|---|---|
| Average handle time | Agent efficiency | Case difficulty | Remove it from individual targets, keep it as a mix diagnostic |
| CSAT | Service quality | Mostly case mix | Compare within intent, never in aggregate across the change |
| Deflection rate | Automation coverage | Little on its own | Always pair with the 7-day repeat contact rate |
| Occupancy | Capacity utilisation | Burnout exposure | Lower the target, then re-derive it from observed data |
Handle time going up is a pass, not a fail
If your residual handle time did not rise after deflection, one of two things is true. Either the automation is taking hard cases it should not be taking, or it is closing conversations that were never resolved. Both are worth investigating, and neither is good news.
Treat a flat line as the anomaly and go looking for the reason.
Aggregate CSAT deserves the same scepticism. The contacts that left your human queue were the ones customers rated most highly, because they were quick and they worked. Removing them lowers the human average without any agent doing worse. Compare like with like or you will performance-manage people for a change you made yourself, which is a specific case of the wider problem in rewriting performance criteria once AI sits inside the work.
The apprenticeship problem nobody budgeted for
The post-deflection queue needs experienced agents, because everything left in it requires judgement. The uncomfortable part is that deflection removes the mechanism that produced experienced agents.
Tier one was never only a cost centre. It was where new hires learned the product, the systems, the tone and the edge cases, on contacts where being slow or wrong was cheap. Automate it and you have kept the requirement and deleted the training ground.
Brynjolfsson, Li and Raymond studied the staggered rollout of a conversational assistant across 5,179 customer support agents and found a 14% average gain in issues resolved per hour. The average conceals the finding that matters: 34% for novice and low-skilled workers, and minimal effect for experienced and highly skilled ones.
The assistant is a substitute for inexperience. It is not a substitute for expertise, and expertise is what the residual queue consumes.
The entry-level cohort is already thinner
This is not a forecast. It is visible in payroll data now. Stanford's Digital Economy Lab, using ADP administrative records through June 2026, reports that workers aged 22 to 25 in the most AI-exposed occupations show a 19% employment decline relative to less-exposed peers. The same paper finds no evidence of economy-wide displacement, and finds the decline is driven by reduced hiring rather than higher departures.
Customer service representatives sit squarely in the exposed group. The Bureau of Labor Statistics counts 2,666,000 of them in 2025 and projects a 5% decline through 2035. It still expects about 289,500 openings a year, all of them from replacement rather than growth. The occupation is shrinking and churning at the same time.
That combination is the trap. You will keep needing to fill seats, and the pool of people who learned the job on easy tickets will keep getting smaller. The identical dynamic is playing out one department over, in the breaking pipeline for junior developers, and support leaders have the advantage of watching it happen there first.
If I were running this, I would fund a deliberate training pipeline out of the deflection savings, in the same budget line, in the same quarter. Route a fixed share of automatable contacts to new hires on purpose.
What the public record shows about cutting too far
Three organisations have now run this experiment in public, and the sequence is consistent enough to plan against.
Klarna deployed an OpenAI-based assistant in February 2024. It handled 2.3 million conversations in the first month, work the company described as equivalent to 700 full-time agents, and it resolved cases in about 2 minutes against 11 for humans. In May 2025 chief executive Sebastian Siemiatkowski told Bloomberg the company had gone too far. The published quote is that investing in the quality of human support is the way of the future for the company. Klarna began recruiting people again, while the assistant still handled about two thirds of inquiries.
Salesforce is the more instructive case, because it did not reverse. Marc Benioff confirmed on the Q2 FY2026 earnings call that customer support headcount had gone from about 9,000 to about 5,000, with AI agents handling 1.5 million inquiries over 9 months. He also said plainly that a great deal cannot be resolved by the agent and has to be escalated to humans, and the company reported an even split between AI-handled and human-handled interactions. The vendor comparison behind those deployments is broken down in the review of Agentforce, Fin and Breeze as support agents.
Gartner has moved in the same direction across four releases. In March 2025 it forecast that agentic AI would autonomously resolve 80% of common customer service issues by 2029. By September 2025 it was predicting that no Fortune 500 company would have fully eliminated human customer service by 2028. In February 2026 it added that half of the companies attributing cuts to AI will rehire for similar functions under different job titles by 2027.
I think that rehiring is mostly a repricing rather than a reversal. The roles come back at a higher skill level, with different titles, and often at higher pay, because the work that remains is genuinely harder. That is a cost outcome very different from the original business case. The function-by-function version of that sum is in the analysis of where agent payback actually lands.
Where this argument is weakest
Three things here would not survive a determined challenge.
No public dataset measures the residual
The core claim, that residual handle time rises after deflection, is a mechanism argument rather than a measured one. I could not find a single credible public dataset that reports pre-deflection and post-deflection handle time for the same queue, split by intent.
The arithmetic in this post is therefore illustrative, and labelled as such in the figure and the table. The premise, that automated intents are faster than average, is almost certainly true in most queues, and you should confirm it in yours rather than take it from here.
Salesforce's own numbers cut the other way
Salesforce reported that its remaining human agents handled roughly the same number of conversations at nearly identical satisfaction scores after the cut. If the residual queue were meaningfully harder per case, you would expect one of those two numbers to move.
Several explanations fit. The routing may be sending more than tier-one work to the agent. The team may have absorbed the difficulty without it showing up in either metric yet. But the honest reading is that a large, well-instrumented deployment has published figures that do not support the concentration story, and one counterexample from a company with real data outweighs a lot of reasoning.
The third weakness is scale. Gartner's own survey found only 20% of service leaders have actually cut agent headcount because of AI. A post about what happens at 60% deflection is addressing a minority of organisations. Most teams are approaching that situation rather than living in it.
Frequently asked questions
What is a good support deflection rate in 2026?
There is no credible public benchmark. The percentages circulating on this question come from vendor blogs and aggregator posts whose figures disagree by wide margins, and they define deflection inconsistently. Measure true resolution instead, meaning the share of conversations that end without a human and without the customer returning within 7 days. That number is comparable across your own quarters, which is the only comparison you need.
Does AI deflection reduce support headcount proportionally?
No, and assuming it does is the most common planning error. Automation clears the intents that are easiest to specify, and those are typically the shortest. Deflecting 60% of contacts that averaged 5 minutes against a queue average of 8 removes about 37% of the handle minutes, not 60%. Build capacity plans on total handle minutes rather than contact counts, and the gap disappears from your forecast.
Why does average handle time go up after AI deflection?
Because the case mix changed, not because agents got slower. The contacts that automation absorbs are short, well-defined and repetitive. What remains is ambiguous, cross-system, high-consequence or already broken by a failed automated attempt. A rise in residual handle time is the expected outcome. If handle time stays flat, check whether the automation is closing conversations that were never actually resolved.
How should you restructure a support team after AI deflection?
Move capacity planning from contacts to minutes, lower the occupancy target while the residual mix settles, and collapse tier one and tier two into a single diagnostic role with real system authority. Then staff a named owner for contacts that arrive after an automated failure, because that stream is new and carries the only reliable signal about where the automation is breaking.
Are companies rehiring customer service staff they cut for AI?
Gartner predicts that by 2027, half the companies that attributed headcount reduction to AI will rehire staff for similar functions under different job titles. Klarna publicly reversed in May 2025 after quality complaints and resumed recruiting human agents. The pattern is less a reversal than a repricing, since roles tend to return at a higher skill level and a higher cost than the ones removed.
What happens to entry-level support jobs when AI handles tier one?
They shrink, and with them the training path to senior roles. Stanford's Digital Economy Lab finds workers aged 22 to 25 in AI-exposed occupations down 19% relative to less-exposed peers, driven by reduced hiring rather than departures. The Bureau of Labor Statistics projects a 5% decline in customer service representative employment through 2035 while still expecting about 289,500 replacement openings a year.
Where to start this week
One query and one decision, in that order.
The query: pull handle time by intent for the 90 days before your automation went live, then split those intents into the ones the agent now handles and the ones it does not. Multiply each group by its volume. You now have the real minutes removed and the real minutes remaining, and your capacity plan has a denominator that reflects the work. If you have not deployed yet, run the same query anyway.
The decision: pick the share of automatable contacts you will deliberately route to new hires, and write it into next quarter's budget as a training cost rather than an efficiency loss. The number can be small. What matters is that somebody has decided it on purpose, because the alternative is discovering in 18 months that you have a queue of hard cases and nobody who learned the job.
Related on this site
The vendor side of this question is covered in the comparison of Agentforce, Fin and Breeze. The escalation design that decides how much reaches a person is in the piece on human-in-the-loop as architecture.
References
- Gartner, Half of companies that cut customer service staff due to AI will rehire by 2027, 3 February 2026. Used for the 50% rehiring forecast and the October 2025 survey of 321 service leaders.
- Gartner via Businesswire, None of the Fortune 500 will have fully eliminated human customer service by 2028, 10 September 2025. Used for the Fortune 500 prediction and Kathy Ross's comments.
- Erik Brynjolfsson, Danielle Li and Lindsey Raymond, Generative AI at Work, Quarterly Journal of Economics 140(2), 2025. Used for the 5,179 agent sample and the 14%, 34% and minimal-effect findings.
- Erik Brynjolfsson, Bharat Chandar and Ruyu Chen, Canaries in the Coal Mine, Stanford Digital Economy Lab, revised 12 August 2026. Used for the 19% relative decline among workers aged 22 to 25.
- US Bureau of Labor Statistics, Customer Service Representatives, Occupational Outlook Handbook. Used for 2025 employment, the 2025 to 2035 projection and annual openings.
- CNBC, Salesforce CEO confirms 4,000 layoffs because I need less heads with AI, 2 September 2025. Used for the 9,000 to 5,000 support headcount figure.
- CX Dive, Salesforce still sees a place for live customer service agents after massive cuts, 4 September 2025. Used for the escalation comment and the 1.5 million inquiries figure.
- CX Dive, Klarna changes its AI tune and again recruits humans for customer service, 9 May 2025. Used for all Klarna figures and the Siemiatkowski quote, sourced from a Bloomberg interview.
Weakest thing about this source base: the central claim that residual handle time rises after deflection has no measured public dataset behind it, and the arithmetic used to illustrate it is constructed rather than observed. Two Gartner figures come from press release text the newsroom would not serve to an automated fetch, taken as reported rather than from the page itself. The Klarna quotes are secondary coverage of a Bloomberg interview.
Related reading