From Sanskriti Khandelwal | Product & Market Analysis

Middle Management After Agents: Coordination Automates, Coaching Does Not

On this page

The average manager now carries 12.1 direct reports, up from 10.9 a year earlier. Agents are good at the part of middle management that moves information around, and poor at the part that decides, coaches and carries blame. Flattening trades a coordination cost you can see for a coaching deficit you cannot. Almost nobody is pricing the second one.

Key takeaways

  • Span of control rose faster in 2025 than in any recent year. Gallup puts the average at 12.1 direct reports, against 10.9 in 2024 and roughly half as many again as 2013.
  • The automatable part of the job is the larger part, and it is the cheaper part. Status collection, scheduling and escalation routing are information work. Judgement, feedback and accountability are not.
  • Manager engagement fell nine points in three years, from 31% to 22%. The layer being asked to absorb wider spans is the layer already coming apart fastest.
  • Manager support is the single strongest lever on team engagement in Gallup's 2026 data. Employees with an AI plan, weekly use and a supportive manager hit 53% engagement, against 31% overall.
12.1Average direct reports per manager in 2025, up from 10.9 in 2024. Source: Gallup, January 2026.
22%Manager engagement worldwide in 2025, down from 31% in 2022. Source: Gallup, 2026.
20%Share of organisations Gartner predicts will use AI to flatten structures through 2026, cutting more than half of middle management roles. Source: SHRM on Gartner.

What agents actually take off a manager's desk

Ask what a middle manager does and you get two very different lists. One is a list of information movements: collecting status, writing it up, chasing a blocker, booking the meeting where the blocker gets discussed. The other is a list of judgements: who is struggling, which of two defensible options to pick, whose promotion to argue for.

Agents are already competent at the first list and largely absent from the second. That split is the whole story, and it is more useful than any headcount forecast.

The routing layer, not the deciding layer

Most of the coordination work a manager does is reading from systems of record and writing a summary somewhere else. A ticket tracker, a CRM, a deployment log, a spreadsheet of last week's numbers. An agent with read access to those systems can produce the same summary continuously rather than weekly.

That is not a small change in degree. Weekly status reporting exists because gathering the information was expensive. When gathering becomes close to free, the artefact loses its reason to exist. The meeting that consumed the artefact usually goes with it.

The deciding layer is untouched by this. Knowing that two teams have collided on a dependency is information work. Choosing which one gets to keep its schedule is not, because the inputs are political, partial and contested. Agents summarise the collision beautifully and have nothing to say about who should lose.

Status reporting is the first thing to go because it is the task with the clearest ground truth. The systems already hold the answer. Nobody disputes what a build log says. There is no ambiguity to resolve and no relationship to damage.

Compare that with a decision on headcount, where the source data is a set of opinions held by people with incentives. This is the same distinction that separates agent pilots that reach production from the ones that stall, examined in more detail in the analysis of where agent pilots actually fail. Ground truth availability predicts automatability better than task complexity does.

The coordination half of the job is large, and it is measurable

If the automatable part were 10% of the role, none of this would matter. The available evidence says it is much larger, though the evidence is thinner than the confidence around it.

16.5 hours, and the paradox inside it

The vendor In Parallel surveyed 247 managers across five countries and reported the result in July 2026. Managers spent about 16.5 hours a week on coordination: 8 hours in status meetings, nearly 5 re-explaining context and nearly 4 searching for information that already existed. That is a vendor survey with a small sample, so treat the specific hours as directional rather than settled.

The interesting finding is the one that cuts against the vendor's own interest. Managers who used AI daily reported the heaviest coordination load, around 20 hours a week, against 13 for occasional users. Ninety-one per cent said they had to load context manually before an assistant became useful.

Two readings are available and only one is comfortable. Either heavy AI users are managers whose jobs were already coordination-dense, or current tools add a context-feeding tax that cancels the saving. I lean toward the second, because the same survey found only 9% saying their assistant usually knew enough to be useful without setup.

Where a manager's coordination week goes Self-reported hours per week. Survey of 247 managers, five countries, July 2026. Status meetings 8 hrs Re-explaining context nearly 5 hrs Searching for data nearly 4 hrs About 16.5 hours in total. Managers using AI daily reported the heaviest load, near 20 hours, against 13 for occasional users. Vendor survey, small sample. Directional, not measured across a population.
Notice the direction of the AI finding. Heavier tool use came with more coordination, not less, which is the opposite of what the category promises.

Even at face value, 16.5 hours is the size of the target, not the size of the saving. Some status meetings exist to make a decision, not to share an update. Some context re-explanation is how a team actually aligns, and removing it would cost more than it saved.

The honest version of the claim is narrower. A large share of a manager's week is spent moving information that a system already holds, and that share is now technically addressable. What gets recovered depends on whether the organisation removes the meeting or just adds a summary to it.

Span of control moved before the technology did

The most quoted number in this debate is a Gartner prediction. The most useful one is a measurement, and it says the flattening started before agents were deployable at scale.

Gallup's January 2026 analysis puts the average number of direct reports per manager at 12.1 in 2025, up from 10.9 in 2024 and nearly 50% above 2013. Amazon set the tone for this in September 2024, when Andy Jassy told staff the company would increase the ratio of individual contributors to managers by at least 15% by the end of Q1 2025.

Read Jassy's reasoning rather than the ratio. He described pre-meetings for the pre-meetings, a longer line of managers, and decisions drifting away from the front line. That is a bureaucracy argument, not an automation argument, and it predates any credible agent deployment.

This matters for how you interpret the current layoff wave. A good deal of what is being attributed to AI is a correction of 2021 and 2022 hiring. That pattern is taken apart in the piece on how much of the layoff wave is an overhiring correction. Agents arrived in time to be given credit for a decision that was already being made.

Span of control was already climbing before agents shipped Average direct reports per manager. Gallup, January 2026. 13 10 7 8.1 10.9 12.1 2013 2024 2025 The 2013 point is derived from Gallup's statement that 2025 is nearly 50% above that year. The median team is still 5 to 6 people. The mean is pulled up by a minority of very wide spans.
The single-year jump is the striking part. One year moved the average almost as far as the previous decade did.

What does not automate, and why the reason is structural

The tempting conclusion is that whatever survives automation survives because agents are not yet clever enough. I think that is wrong. The residue has a specific shape, and the shape is not about capability.

Judgement when the data runs out

Managers are paid to decide when the evidence is incomplete and stays incomplete. Which of two competent people gets the stretch project. Whether a slipping deadline is a resourcing problem or a competence problem. Whether to escalate now or give it another week.

An agent can lay out the considerations. It cannot own the choice, because owning a choice means being answerable for it later. Accountability is a social property, not a computational one, and it does not transfer to a system that cannot be demoted.

Coaching, and the evidence that it still pays

The commercial case for coaching is stronger than the sentimental one. Gallup's July 2026 analysis of 43,262 responses found engagement at 48% among employees with manager support on AI, against 30% without it, with US engagement overall stuck at 31%. Employees with all three conditions, an organisational plan, weekly use and a supportive manager, reached 53%.

Gallup's span work points the same way. Employees who strongly agreed they received meaningful feedback each week were engaged at roughly 7 in 10, regardless of how large their team was. Among those who disagreed, the figure was 1 in 4.

That is the finding I would put on the wall. Team size is not what determines whether management works. Feedback frequency is. And feedback frequency is exactly the thing that a rising span quietly destroys, because it is the only part of the job with no economies of scale.

Weekly coaching minutes per report, at a fixed budget of four hours.
Direct reportsMinutes per report per weekWhat that buys
548 minutesA real conversation, with follow-up.
830 minutesA proper check-in, no follow-up.
1220 minutesA status exchange with a coaching label on it.
1516 minutesA queue.
25About 10 minutesNothing anyone would call development.

This table is arithmetic on an assumption, not measured data. The four hour weekly coaching budget is illustrative. Change the budget and the ranking does not change, because the constraint is division, not effort.

Flatter is not free, and the bill arrives late

The case for removing a layer is easy to state and easy to measure. The costs are real, slower to appear, and they land in places the reorg spreadsheet does not have a column for.

The layer absorbing the change is the layer disengaging fastest

Gallup's State of the Global Workplace 2026 reports manager engagement at 22% in 2025, down from 27% in 2024 and 31% in 2022. Non-managers went from 20% to 19% across the same window. The premium that used to come with the job has almost gone.

Set that beside a rising span and the sequencing looks poor. Organisations are widening the responsibilities of the group whose engagement is falling fastest, and expecting added coordination capacity from agents to cover the difference.

Coordination disperses, it does not vanish

Removing a layer does not remove the coordination that layer performed. It redistributes it to people who now do it part-time, without the context, and usually without the authority to settle anything.

The failure mode is familiar to anyone who has run an autonomous-team model at scale. Duplicate work appears, systems diverge, and dependencies surface late because no single person was watching the seams. This is the organisational version of the argument for keeping a person in the loop at defined checkpoints, laid out in the piece on where a human review step actually belongs.

Three ways flat structures fail, and what each looks like early.
FailureMechanismEarly signal
Coordination returns as meetingsWork that a layer absorbed becomes everyone's part-time jobCalendar load rises for individual contributors, not managers.
Decisions stallAuthority was not moved down with the workMore items marked blocked, fewer marked rejected.
The bench emptiesNobody is being developed into the next roleExternal hires for every senior opening.

The promotion ladder loses a rung

The manager layer is where people learn to run things at low stakes before running things at high stakes. Delete it and you have removed the training environment for your own future leadership, which is a cost that shows up three years later as an external search fee.

This is the same structural problem appearing one level down in engineering, where agents absorb the work juniors used to learn on. That version is traced in the analysis of the breaking junior developer pipeline, and the mechanism is identical: automate the apprenticeship and you get a seniority gap you cannot hire your way out of.

Where this argument is weakest

Two things could make most of the above wrong, and one of them is an official statistic pointing the other way.

The official numbers say the layer is growing

The US Bureau of Labor Statistics projects the opposite. Its January 2026 projections article has management occupations rising from 13.6 million in 2024 to 14.4 million in 2034. That is a gain of 6.1% and the fifth fastest of 22 occupational groups. It is a projection rather than a result, and projections lag structural change. Even so, it is an inconvenient number for anyone declaring the layer finished.

The layoff coverage and the employment data can both be right. Concentrated cuts at large technology firms are highly visible, and they are a small share of a labour market where management jobs are growing in health, construction and operations. Be careful about generalising from Seattle to the economy.

Coaching may automate further than I have allowed

The Conference Board reported in October 2025 that AI could handle roughly 90% of day-to-day career coaching needs. Ninety-six per cent of workers said the responses were tailored to their goals, and 89% got specific next steps. The reported weak points were scripted language, thin contextual memory and no personal connection.

Two of those three are engineering problems that are being worked on now. If contextual memory improves, the boundary I have drawn moves, and the coaching argument weakens considerably. My position is that the boundary holds for feedback that carries consequences, because those conversations derive their weight from the fact that a person with power said them. I hold that position less firmly than the rest of this piece.

What the layer actually becomes

My reading is that middle management gets smaller in headcount and harder in content. The routing work leaves, the judgement work stays, and the judgement work does not get easier when it is concentrated into fewer people.

The job description changes in a specific direction. Less reporting upward, more deciding downward. Less chasing status, more setting the thresholds at which an agent should stop and ask a human. Someone has to define what a system is allowed to decide alone, and that is a management task before it is a technical one.

Two adjacent shifts follow. Performance criteria have to be rewritten when the output is a joint product of a person and a system, a problem worked through in the piece on performance review criteria in an agent workplace. And the payback varies enormously by function, which is why a uniform span target is a mistake, as the function-by-function view in the breakdown of agent payback by function shows.

The unbundling: which parts of the job move, and what breaks if nobody owns them.
TaskWhere it landsWhat breaks if it is dropped
Status collection and reportingAgent, reading systems of recordNothing. This was always a tax.
Scheduling and meeting logisticsAgentNothing.
Escalation routingAgent, against a human-set thresholdEscalations reach the wrong person late.
Prioritising a contested backlogHuman, informed by agent summariesThe loudest stakeholder wins by default.
Weekly feedback to each reportHuman. An agent can draft, not deliver.Engagement falls exactly as span rises.
Deciding on incomplete evidenceHuman, on the recordThe decision waits for data that never arrives.
Performance calls and exitsHuman, with a named ownerLegal exposure and a collapse in trust.
How far an agent can take each part of the job Assessment based on whether the task has stable ground truth and whether a person must be answerable for it. AGENT DOES IT AGENT ASSISTS HUMAN ONLY Status reporting Meeting scheduling Escalation routing Resource forecasting Prioritising a backlog Weekly feedback Promotion decisions Exiting someone The dividing line is accountability, not difficulty. Forecasting is technically harder than firing someone.
The bottom three rows are not waiting on better models. They require someone who can be held responsible afterwards.

How to redesign the role before a reorg does it for you

Waiting for a restructure to define the job is the worst available option. A restructure optimises for a headcount number, not for what the layer is for.

Start by separating the two lists explicitly. Write down every recurring thing a manager on your team does, and mark each one as information movement or judgement. The information movement items are candidates for automation now. The judgement items are the job description you should be hiring and promoting against.

Then set a span ceiling on the judgement half, not on the whole role. If you expect a manager to give meaningful weekly feedback, and you believe that takes 30 minutes per person, you have set a maximum span whether you admit it or not. Firms reporting unusually high output per head have generally made this trade deliberately, a pattern visible in the numbers on revenue per employee at AI-native firms.

Finally, be sceptical of hour-saving claims made about your own calendar. The recovered time gets consumed by something, and if nobody names what, it will be consumed by more meetings. The same accounting problem shows up in the hours math behind the four-day week claims.

Frequently asked questions

Will AI replace middle managers?

Not wholesale, but it is unbundling the role. Agents handle status collection, scheduling and escalation routing well, because those tasks have stable ground truth in existing systems. They do not handle judgement calls, weekly feedback or accountability, because those require someone answerable afterwards. The realistic outcome is fewer managers doing a harder, more decision-heavy job, rather than the layer disappearing.

What percentage of a manager's job can AI automate?

No reliable measured figure exists. A vendor survey of 247 managers reported about 16.5 hours a week spent on coordination, which is the size of the target rather than the size of the saving. Some of those hours are decision meetings wearing a status label. Treat any single percentage you see as an estimate built on a small sample, not a measurement.

What is a good span of control in 2026?

Gallup's data shows an average of 12.1 direct reports in 2025, with the median team still at 5 to 6 people. The more useful test is feedback capacity. Employees receiving meaningful weekly feedback were engaged at roughly 7 in 10 regardless of team size, so set your span from how much coaching time you can genuinely fund per person.

Why are companies cutting middle management?

Mostly for reasons that predate agents. Amazon's Andy Jassy framed his 2024 push to raise the individual contributor to manager ratio as an attack on bureaucracy and slow decisions, not on cost. A large share of current cuts is also a correction of pandemic-era overhiring. AI is a genuine contributor and a convenient explanation for decisions already under way.

Can AI do performance reviews and coaching?

Partly. The Conference Board found in October 2025 that AI could cover roughly 90% of daily career coaching needs, with workers rating the responses as tailored and actionable. The reported weaknesses were scripted language, weak contextual memory and no personal connection. Consequential conversations, including promotions and exits, still need a person who can be held responsible for the outcome.

What should a middle manager do to stay relevant?

Move deliberately toward the judgement half of the job. Automate your own status reporting before someone does it for you, then reinvest the time into weekly feedback that people would notice if it stopped. Learn to set the thresholds at which an agent must stop and ask a human, because defining that boundary is becoming a core management task rather than a technical one.

Where to start this week

Two things, both small enough to finish before Friday.

First, audit one manager's calendar for a single week and tag every entry as information movement or judgement. You will get a ratio specific to your organisation, which is worth more than any survey figure in this post, including the ones I have cited.

Second, calculate your real coaching minutes per report. Take the time a manager actually protects for one-to-ones, divide by span, and compare the result against the table above. If the answer is under 20 minutes, your span is already past the point where the non-automatable half of the job can be done at all.

Related on this site

If you are redesigning roles around agents, the two adjacent pieces are performance criteria when a system does part of the work and where agent payback actually shows up by function.

References

  1. Gallup, Span of Control: What's the Optimal Team Size for Managers?, 13 January 2026. Used for span of control figures and the weekly feedback finding.
  2. Gallup, State of the Global Workplace 2026, 2026. Used for manager engagement of 22% in 2025 against 31% in 2022.
  3. Gallup, Employee Engagement Remains Flat as AI Adoption Accelerates, 21 July 2026. Used for the manager support and engagement figures.
  4. US Bureau of Labor Statistics, Industry and occupational employment projections, 2024 to 2034, Monthly Labor Review, January 2026. Used for management occupation projections.
  5. Amazon, Update from Andy Jassy on return to office plans and manager team ratio, September 2024. Used for the 15% ratio target and the bureaucracy reasoning.
  6. SHRM, Transforming Work: Gartner's AI Predictions Through 2029. Used for the Gartner flattening prediction attributed to Daryl Plummer.
  7. No Jitter, Managers' busywork increases with AI use, July 2026, reporting In Parallel's survey of 247 managers. Used for all coordination hour figures.
  8. HR Dive, AI can provide most career coaching, but human expertise is still needed, 22 October 2025, reporting The Conference Board. Used for the coaching automation figures.

The weakest source here is reference 7. It is vendor research with a sample of 247 self-selected managers, and every coordination hour figure in this post rests on it alone. The Gallup and BLS material is far stronger, and where the two disagree, prefer the larger sample.

SK
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading