From Aryan Vatsa | Product & Market Analysis

AI Chief of Staff: The Meeting Work Agents Take, and the Work They Cannot

On this page

Meeting agents are winning every part of the job you can check against a transcript. In a randomised trial of more than 6,000 workers, Microsoft 365 Copilot cut email reading time by 30 minutes a week and moved total meeting time not at all. That gap is the entire map of what an AI chief of staff can take from you, and what it cannot.

Key takeaways

  • Scheduling, capture and follow-up transfer. Judgement does not. The dividing line is not task difficulty. It is whether the output can be checked against a written record by someone who was not in the room.
  • The measured gains land on solo work. Across 56 firms, regular Copilot users read email 30 minutes less each week and completed documents 12% faster, while total meeting time showed statistically insignificant change.
  • Agents do worse on administrative work than on engineering. On a 175-task benchmark built inside a simulated company, the strongest agent completed 30.3% of tasks autonomously, and admin tasks scored below software development.
  • The 400-to-1 cost ratio is the wrong argument. A Copilot seat lists at $18 to $30 per user per month against mean chief of staff total compensation of $192,100. The two do not cover the same work, so the ratio proves nothing.
30.3%Best agent's autonomous completion rate across 175 real professional tasks. Source: TheAgentCompany, NeurIPS 2025.
30 minWeekly email reading time saved by regular Copilot users, against no significant change in meeting time. Source: Microsoft Research, April 2025.
57%Share of meetings that happen ad hoc, with no calendar invite at all. Source: Microsoft Work Trend Index, 2025.

What an AI meeting agent actually does in 2026

An AI meeting agent joins a call, transcribes it, drafts a recap, extracts action items and pushes them into a task tool. Some also book the meeting and chase the follow-up. It does not decide who should be in the room, what a decision costs politically, or who needs to hear the news first.

That is a narrow product. It is also a genuinely useful one, and the narrowness is the interesting part.

The four jobs on the calendar

Strip the marketing away and the category does four things. It finds a slot across many calendars. It captures what was said. It drafts a summary with owners and dates. It chases those owners afterwards.

Each of those has a property the rest of the executive support job does not have. You can tell within minutes whether the agent got it right, by opening the recording or the calendar.

What Microsoft shipped, and what it held back

Microsoft's Facilitator agent in Teams reached general availability in August 2026. It offers shared live notes, an agenda with a timer, in-meeting question answering, task tracking and a recap. It requires a Microsoft 365 Copilot licence and works only in scheduled Teams meetings.

Two details in the rollout notes are worth more than the feature list. Facilitator is off by default, and it triggers less than once per meeting on average. A vendor that believed its agent improved live discussion would not ration the interventions that carefully.

Read that as a confession rather than a limitation. The parts of a meeting an agent can safely touch are the parts around the edges, before it starts and after it ends.

The transfer test that predicts which tasks move

Most coverage of this question sorts tasks into routine and strategic. That split does not survive contact with the data, because agents fail at plenty of routine work and succeed at some complicated work.

A better test asks one question. Can a person who was not present verify the output against a record?

Checkable against the record

Booking a slot is checkable against five calendars. A transcript is checkable against audio. An action list is checkable against the transcript. A chase message is checkable against the action list.

Every one of these has a ground truth that exists independently of the agent. When an agent is wrong, you find out immediately, and the cost of being wrong is one correction.

Checkable only against the room

Now take the other half of the job. Deciding that the pricing conversation happens after the headcount conversation, not before. Deciding that a director hears about a reorganisation from you rather than from the deck. Deciding which of two loud people in a review is actually right.

None of these has a record to check against at the moment the call is made. The verification arrives weeks later, in how people behave. An agent optimising against a reward it cannot observe is not doing the task, it is guessing at it.

This is the same boundary that shows up in the architecture of human review in agent systems. Where an error is cheap and visible, automate. Where an error is expensive and invisible for a quarter, do not.

Where the time actually moved, measured twice

Two studies published since 2025 make this concrete. One measures what happened in real companies. The other measures what agents can do when nothing stands in their way.

The 6,000-worker experiment

Microsoft Research ran a randomised controlled trial of Microsoft 365 Copilot across more than 6,000 workers at 56 firms, allocating licences at random and watching for six months. It is the largest clean read anyone has published on this software.

The results split neatly along the transfer test. Licensees read email 12 minutes less each week, a 7% fall. Regular users read email 30 minutes less, an 18% fall. Word documents were completed 6% faster for licensees and 12% faster for regular users, against an average document taking 7 days.

Then the meetings. Total meeting time showed small and statistically insignificant increases. Workers joined meetings 5 minutes later and left 10 minutes earlier, so the shape of attendance moved. The total cost did not. Firm-level effects ranged from a 26-minute weekly increase to a 25-minute decrease, a spread wide enough to caution against any single headline.

Where a Copilot licence moved the needle, and where it did not Randomised trial, 6,000+ workers at 56 firms, six months. Effect on regular users. no change. Email reading time. -18% 30 minutes back each week. Document completion time. -12% roughly one day off a 7-day document. Total meeting time. not significant. joined 5 min later, left 10 min earlier. Solo work compressed. Coordinated work did not.
The two bars pointing left are tasks you do alone. The stub on the right is the task that needs other people, and it did not move.

The benchmark that ranks admin below engineering

The second study is TheAgentCompany, built by Carnegie Mellon researchers and published at NeurIPS in 2025. It puts agents inside a simulated software company with internal websites, files and colleagues. It then gives them 175 long-horizon professional tasks across software development, project management, data science, administration, human resources and finance.

The strongest agent, Gemini 2.5 Pro, completed 30.3% of tasks autonomously, and 39.3% with partial credit. Claude 3.7 Sonnet reached 26.3% and 36.4%.

The finding that matters here is not the headline rate. It is the ordering. Software engineering tasks scored higher than the seemingly simpler administrative and finance tasks. The category everyone assumed would fall first held up better than the category everyone assumed was already automated.

The authors name the reason. Agents failed to understand conversational implications, asked appropriate questions and then failed to act on the answers, and did not recognise when a task was socially incomplete. They also took shortcuts under difficulty, renaming an entity rather than finding the correct one.

That last behaviour should alarm anyone planning to hand an agent a coordination role. A chief of staff who quietly renamed a problem instead of solving it would be fired. It is a specific instance of the pattern catalogued in the review of how agent pilots actually fail.

The task transfer map

Put the two studies together with the product reality and the map is not ambiguous. Here is the executive support job, sorted by the transfer test rather than by seniority.

Which chief of staff tasks transfer to a meeting agent, and what decides it
TaskTransfers todayWhat decides it
Finding a slot across many calendarsYesGround truth is the calendar. An error is visible instantly.
Transcribing and summarising a callYesGround truth is the recording. Errors are correctable in one pass.
Extracting action items with ownersMostlyCheckable against the transcript. Ambiguous ownership still needs a human.
Chasing an owner for statusPartlyThe message is checkable. Knowing when chasing will backfire is not.
Setting the agenda and its orderNoSequencing is a political decision with no record to verify against.
Deciding who is in the roomNoInclusion signals status. The cost of a wrong call is invisible for weeks.
Delivering an unwelcome decision to a peerNoRequires accountability. Nobody accepts a refusal from a scheduling bot.

Verdicts are the author's, applying the checkability test to the role description published by Harvard Business Review. They are a framework, not a measurement, and the first four rows will keep improving.

The same seven tasks, scored on two properties Green transfers. Amber transfers with supervision. Red does not. Output checkable. Needs the room. Transfers. Calendar slot finding. Yes No Yes Transcript and recap. Yes No Yes Action item extraction. Yes Some Mostly Chasing an owner. Yes Yes Partly Agenda and its order. No Yes No Who is invited. No Yes No Delivering a refusal. No Yes No Read the middle column. Every red verdict needs the room, not the record.
The pattern is a single column deep. Nothing that needs the room transfers, and nothing that avoids the room stays human for long.

What does not transfer, and why capability will not fix it

The comfortable version of this argument says the hard half is safe because models are not good enough yet. That is not my position. The hard half is different in kind, not in difficulty.

Sequencing is a political act

Consider the smallest possible chief of staff decision. Two items on one agenda, and you choose which goes first. There is no correct answer available from the transcript. The right order depends on who is in a bad mood, which decision constrains the other, and whether one attendee needs a win before absorbing a loss.

An agent can be told those constraints. It cannot observe them, because nobody writes them down, and the people who hold them would not put them in a document if asked. The same problem sits underneath the question of what authority you can contractually grant an agent.

The context nobody wrote down

Microsoft's own telemetry makes the scale of this visible. Its 2025 Work Trend Index found that 57% of meetings happen ad hoc, with no calendar invite at all. Workers are interrupted every 2 minutes during core hours, roughly 275 times a day.

An agent that lives on the calendar cannot see the majority of the meetings. The unscheduled conversation in the corridor is where the sequencing information lives, and it produces no artifact.

The executive time data says the same thing from the other end. Porter and Nohria tracked 27 chief executives for 13 weeks and found they worked 62.5 hours a week, spent 72% of work time in meetings, and advanced their own agenda only 43% of the time. A separate diary study of 94 Italian chief executives found only 15% of their time was spent working alone.

Read those two numbers together. The compressible share of an executive week is the 15%, and that is precisely the share the Copilot trial moved.

The executive week an agent is trying to compress 27 chief executives, tracked around the clock for 13 weeks. Average week 62.5 hours. HOW THE TIME IS SPENT. 72% in meetings with other people 28% everything else. WHOSE AGENDA IT SERVES. 43% own agenda 57% set by other people and events. A separate diary study of 94 chief executives put solo working time at 15% of the week.
An agent that only compresses solo work is competing for the smallest block on this chart. The dark blue bar is where the job is.

The cost comparison most teams run backwards

The pitch deck version of this comparison is a seat price against a salary. It produces a ratio of about 400 to 1 and it is close to meaningless.

Microsoft 365 Copilot Business lists at $18 per user per month paid yearly, discounted from $21 through 31 December 2026, with the enterprise add-on at $30. Call it $360 a year at the top of the range.

The Chief of Staff Association reported mean salary of $158,500 and mean total compensation of $192,100 in its 2023 survey. That survey was weighted 71% toward North America and drew on self-selected members, so treat it as directional.

Dividing one by the other tells you nothing, because the agent covers four rows of the table above and the human covers all seven. The useful comparison is narrower. Take the hours your executive support currently spends on capture and chase, price those hours, and set the licence against that number only. That is the same discipline applied in the function-by-function agent payback analysis.

The labour statistics support caution rather than replacement. The US Bureau of Labor Statistics puts executive secretaries and executive administrative assistants at 489,300 jobs in 2025 and 489,200 in 2035, a change of zero, while the broader secretarial category falls 2%. Fortune reported that the wider collapse already happened, from roughly 3.5 million secretarial roles in 2004 to 2.1 million in 2024.

My reading is that the executive-support tier stopped shrinking at the point where the remaining work stopped being checkable. Word processing and diary management were automated over two decades. What is left is the part the transfer test says agents cannot take.

Where this argument is weakest

Four places, and one of them may eventually undo the whole framework.

A simulation is not an office

TheAgentCompany runs inside a synthetic company with simulated colleagues. That design makes it reproducible. It also makes social failure easier to produce than it would be with real, forgiving humans who repeat themselves. The 30.3% figure is a snapshot on models available at evaluation time, and newer systems would score higher.

The Copilot trial has a matching limit. Six months on early software, with 38% weekly adoption and a firm-level spread from 9% to 75%. A null result on meeting time may simply mean the study ended too soon.

The strongest objection is that my transfer test decays as agent memory improves. Once a system carries a year of your correspondence, some of the context I called unwritten becomes written. I still think sequencing resists this, because that information was never recorded in any channel. But it is an argument I could lose within two product cycles.

Finally, the transcription risk is smaller than the headlines suggest. Researchers found roughly 1% of audio segments produced fabricated text on a 2023 build of Whisper, with about 38% of those fabrications harmful, and a retest after an update removed most of them. That is real, dated, and not a reason to avoid capture. The consent and training-rights exposure covered in the analysis of note-taker consent obligations is the larger practical problem.

Frequently asked questions

Can an AI agent replace a chief of staff?

No, and the gap is not close. Meeting agents handle scheduling, transcription, recap drafting and follow-up chasing, which are the parts of the job with a checkable output. A chief of staff also decides what reaches the executive, sequences unpopular decisions and carries relationship context that was never written down. On a 175-task benchmark of real office work, the strongest agent finished 30.3% of tasks autonomously and did worse on administrative work than on software engineering.

What can AI meeting agents actually do well?

Four things. They find a slot across many calendars, transcribe a call, draft a recap with action items, and chase the owners of those items afterwards. All four produce an artifact you can check against the recording or the calendar within minutes. Microsoft's Facilitator agent in Teams adds a shared live notepad, an agenda timer and in-meeting question answering, though it ships off by default and triggers less than once per meeting on average.

Do AI meeting assistants save time?

They save time on reading and writing, and the evidence on meetings themselves is much weaker. In a randomised trial of more than 6,000 workers at 56 firms, regular Copilot users spent about 30 minutes less each week reading email and finished documents roughly 12% faster. Total meeting time showed small, statistically insignificant increases. Workers joined 5 minutes later and left 10 minutes earlier, so the shape of a meeting moved before its total cost did.

What tasks should you never delegate to an AI agent?

Anything whose correctness depends on knowing the people involved. That covers deciding who attends, what order two contentious items are heard in, how a decision is communicated to the person it disadvantages, and any commitment made on your behalf to a counterparty. These calls have no checkable output at the time you make them. You find out whether they were right months later, from how people behave.

How much does an AI chief of staff cost compared with a human one?

Microsoft 365 Copilot Business lists at $18 per user per month paid yearly, discounted from $21 through 31 December 2026, and the enterprise add-on lists at $30. The Chief of Staff Association reported mean total compensation of $192,100 in its 2023 survey. That is roughly 400 to 1 at list price. The ratio is misleading, because the two do not cover the same work, so compare the agent against hours rather than headcount.

How to redraw the delegation line this quarter

Start with an inventory, not a tool. Take your own calendar for the last two weeks and mark every commitment you made with one letter. C if a colleague could verify you got it right by reading a record. R if only the room could tell.

Then price the C column. Those hours are the entire addressable surface for a meeting agent, and in most executive weeks they are smaller than the vendor assumed. If the C column is under four hours, buy nothing this quarter and revisit in six months.

If it is larger, run one narrow deployment against it and record a baseline first. Count meetings held, minutes in meetings, and days from action item to closure, for four weeks before you switch anything on. Without that baseline you will be reading a vendor dashboard as though it were evidence, which is how deployments end up with negative measured return.

Related on the same question

If the coordination layer is what agents cannot take, the next question is what happens to the people who currently own it. That is the subject of the analysis of middle management after agents.

References

  1. Dillon, Jaffe, Peng and Cambon, Early Impacts of M365 Copilot, Microsoft Research, April 2025. Used for all randomised trial figures on email, documents, meeting time and adoption.
  2. Carnegie Mellon University and collaborators, TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks, NeurIPS 2025. Used for the 175-task completion rates and the failure-mode analysis.
  3. Microsoft WorkLab, Breaking Down the Infinite Workday, Work Trend Index special report, 2025. Used for the ad hoc meeting share and the interruption frequency.
  4. Michael Porter and Nitin Nohria, How CEOs Manage Time, Harvard Business Review, July 2018. Used for the 62.5-hour week, the 72% meeting share and the 43% own-agenda figure.
  5. Bandiera, Hansen, Prat and Sadun, CEO Behavior and Firm Performance, NBER Working Paper 23248. Used for the 15% solo working time in the Italian chief executive diary study.
  6. US Bureau of Labor Statistics, Secretaries and Administrative Assistants, Occupational Outlook Handbook, 2025 edition. Used for the 2025 to 2035 employment projections and median pay.
  7. The Chief of Staff Association, Compensation, State of the Industry Report 2023. Used for mean salary and mean total compensation.
  8. Microsoft, Microsoft 365 Copilot Business pricing, accessed September 2026. Used for all seat prices and the promotional end date.

The weakest source here is the Chief of Staff Association compensation report. It is a self-selected member survey from 2023, weighted 71% toward North America, and it is the only load-bearing figure in this post without a published sample size. The Whisper hallucination rates in the limitation section come from Science coverage of a 2023 model build and are directional only. The secretarial headcount decline from 3.5 million to 2.1 million is as reported by Fortune, July 2026.

AV
Aryan Vatsa
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading