From Sanskriti Khandelwal | Product & Market Analysis
Engineering Burnout in 2026: AI Raised the Pace and Nobody Changed the Team Size
On this page
49% of software engineers now feel emotionally drained at work at least once a week, up from 39% a year earlier. Team sizes did not change. Engineering burnout rose because delivery expectations moved with the tooling and almost nothing else did. That gap is a scheduling problem wearing a wellbeing costume, and it belongs to managers rather than to the people burning out.
Key takeaways
- Exhaustion rose fastest among the people setting the pace. LeadDev's 2026 report puts weekly emotional exhaustion at 49% for engineers and 54% for chief technology officers, against 39% and 24% one year earlier.
- AI adoption did not move the burnout number in either direction. DORA surveyed nearly 5,000 technology professionals, found 90% using AI at work, and found friction and burnout statistically unchanged by that adoption.
- The saved time is real and it does not stay saved. 68% of developers report saving 10 or more hours a week with AI. And 90% report losing 6 or more hours a week to organisational friction the tools never touched.
- The work moved rather than shrank. Generation got faster, review and verification did not. The heaviest AI users report the most evening and weekend work of any group measured.
The gap that opened between pace and capacity
Engineering burnout is rising because AI raised the expected pace of delivery while team size, review capacity and on-call rotas stayed fixed. Faster generation does not shorten review, testing, incident response or coordination. The extra output still has to pass through the same human bottlenecks, and the queue in front of those bottlenecks got longer.
The adoption numbers are no longer in dispute. DORA found 90% of technology professionals using AI at work, at a median of 2 hours per working day. Stack Overflow's 2025 survey of 48,945 developers put usage or planned usage at 84%, up from 76% the year before.
What did not rise alongside it is the number of people available to check the work. This is the part of the story most commentary skips, because it is a resourcing decision rather than a technology one.
My position is that most writing about AI burnout blames the tool, and the tool is the wrong defendant. A code generator has no opinion about your quarterly roadmap. The people who set that roadmap watched output rise and treated the rise as new baseline capacity rather than as a temporary surplus with unpaid downstream costs.
| Stage of delivery | Effect of AI assistance | Who absorbs the change |
|---|---|---|
| Writing a first draft of code | Substantially faster for most respondents | The individual engineer, who feels faster |
| Reviewing that code | Unchanged, and volume per reviewer rises | Senior engineers and tech leads |
| Debugging and verification | Slower for a large minority of developers | Whoever merged it, often after hours |
| Incident response and recovery | Longer average recovery among heavy AI users | The on-call rota, which did not grow |
| Specification and coordination | Heavier, because more capacity needs feeding | Engineering managers and product partners |
| Headcount and budget | Flat or falling across most teams | Nobody, which is the problem |
Rows one to four draw on the DORA 2025 report, the Stack Overflow 2025 survey and the Harness February 2026 survey. Rows five and six are the editorial reading of those findings rather than a measured result. Both are labelled as such.
Burnout is a property of the system, not the workstation
The World Health Organization classifies burn-out in ICD-11 as an occupational phenomenon rather than a medical condition. Its published definition is precise and worth quoting properly: a syndrome resulting from chronic workplace stress that has not been successfully managed.
Three dimensions define it. Energy depletion or exhaustion. Increased mental distance from the job, including cynicism. Reduced professional efficacy. Notice that none of the three is a statement about tooling, and all three are statements about sustained conditions.
That framing predicts the single most useful finding of the past year. DORA looked at AI adoption against eight outcome measures and found higher individual effectiveness, higher delivery throughput and higher delivery instability. Friction and burnout came out at similar levels for high and low adopters once other variables were controlled.
Read that carefully, because it cuts both ways. AI is not making your team burn out. AI is also not going to fix a team that is already burning out, which is what a large number of 2025 rollout business cases quietly assumed.
DORA's own framing is that AI acts as an amplifier of whatever conditions already exist. Strong teams get stronger and struggling teams get more visibly stuck. The report's lowest archetype, the group it calls foundational challenges, combines low performance with high friction and high burnout. Adding tools to that group changes the volume rather than the outcome.
Where the time AI saves actually goes
Atlassian's 2025 developer experience research produced the cleanest description of the problem anyone has published. 99% of developers reported saving time with AI, and 68% reported saving 10 hours or more each week.
The same survey found that 50% lose 10 or more hours a week to organisational inefficiency, and 90% lose at least 6. Finding information, adapting to new tools and switching context between systems account for most of it.
So the hours arrive and the hours leave. If you only instrument the arrival, you will conclude your team has spare capacity. You will be wrong by roughly the same amount every week.
Verification quietly replaced authorship
DORA's qualitative interviews are blunter than its charts. The researchers observed that time saved during initial code generation is often re-allocated to verification overhead, and quoted one developer saying they spend less time writing code and more time babysitting the model.
Stack Overflow's data explains why. 66% of developers named AI solutions that are almost right but not quite as their single biggest frustration, and 45% said debugging AI-generated code takes more time than writing it themselves would have.
Almost-right output is the expensive failure mode. Clearly wrong output is discarded in seconds. Plausible output has to be read line by line by somebody who understands the system, and that reading is cognitively heavier than authorship. That gap between developer usage and developer trust is the daily texture of the work.
Review load lands on the same few people
Generation capacity is distributed across the whole team. Review capacity is not. It sits with a small number of senior engineers who already had the most context and the fullest calendars.
DORA recorded the mechanism directly in an interview: reviewing code is harder than writing it, and AI increases the rate at which people can produce code that needs reviewing. That asymmetry is the actual bottleneck, and it is covered in more depth in the analysis of why code review became the constraint on AI-assisted teams.
There is a second-order cost here that shows up later. When seniors spend their week reviewing rather than teaching, the informal apprenticeship that produced the next generation of seniors stops running. That pattern is examined in the piece on what is happening to the junior developer pipeline.
Why managers keep missing this
Every engineering leader I have spoken to this year has a dashboard showing AI adoption climbing. The same leader has a team that feels worse than it did last year. Both readings are accurate. The instruments are just pointed at different things.
Self-reported speed is not speed
The most uncomfortable result in this field remains METR's randomised trial from July 2025. 16 experienced open-source developers completed 246 real issues on repositories they had maintained for years, with AI use randomly allowed or disallowed per issue.
Developers took 19% longer when AI was allowed. They had forecast a 24% speedup beforehand. After finishing, having just been slowed down, they still estimated that AI had made them 20% faster.
METR is careful about scope and so should you be. 16 developers is a small sample. The repositories averaged more than a million lines of code, and the tools were early 2025 vintage. The authors explicitly decline to claim the result generalises to most software work. What it does establish is narrower but solid. The perception of speed is unreliable in exactly the population managers survey.
Adoption dashboards measure the wrong object
Seat counts, acceptance rates and daily active users tell you a tool is being used. They tell you nothing about whether the work got easier, and a rising line on that chart is frequently read as evidence that it did.
If you want the honest version of this measurement problem, the piece on the three separate scoreboards for AI coding tools lays out which numbers answer which question. The analysis of deployments that returned less than they cost shows what happens when nobody checks.
What actually lowers the load
This is the section the topic deserves, so it is deliberately concrete. None of it requires a budget approval. None of it involves banning tools. That does not work, and it signals distrust.
Cap work in progress before you cap anything else
Faster generation increases the number of things a team can have half-finished at once. Half-finished work is where cognitive load accumulates, because every open thread carries context that has to be held somewhere.
Set an explicit limit on concurrent work items per engineer and enforce it in the board rather than in a meeting. DORA's capabilities model names working in small batches as one of seven practices that decide whether AI amplifies performance or amplifies instability, and small batches are also the cheapest burnout control available to you.
Budget review capacity as explicitly as you budget licences
You would not add a service without provisioning the database behind it. Adding generation capacity without adding review capacity is the same error committed in a domain where the database is a person.
Two practical moves. Give reviewers protected, scheduled review blocks rather than expecting review to happen in gaps. And make review load visible per person, weekly, so the team can see when it is concentrated on two names out of nine. Concentration, not volume, is what turns review into a grind.
Protect the stopping point
Rebecca Koniahgari, a technical lead quoted in LeadDev's reporting, described the mechanism better than any survey has. Every problem has an immediate next step, so the session keeps going until you make a conscious decision to stop.
That is the crux of the always-shipping problem. Traditional development had natural stopping points built in, such as waiting on a build, a review, or another team. Those pauses were annoying and they were also recovery. Agent-assisted work removes them, so the pause has to be reinstated on purpose: time-boxed sessions, a hard end to the working day, and separating exploratory work from execution work. Stopping becomes a deliberate act.
| What to change | What it costs you | What it will not fix |
|---|---|---|
| Hard limit on concurrent work items | Visible queue, and uncomfortable prioritisation conversations | Understaffing. A limit makes it legible, not smaller. |
| Scheduled, protected review blocks | Roughly 4 to 6 hours a week per senior reviewer | Review quality, if the reviewer lacks system context. |
| Time-boxed AI sessions with a defined stop | Some genuinely productive late-evening streaks | Deadline pressure imposed from outside the team. |
| Separate exploration from execution | Extra planning ceremony, which teams resist | Unclear priorities, which is a leadership problem. |
| Rebalance the on-call rota against instability | More people in rotation, or fewer deploys | The instability itself, which needs quality work. |
| Doing nothing and waiting | Nothing now | Nothing. Defensible if your team already scores well on friction. |
What to measure instead of adoption
Harness commissioned Coleman Parkes to survey 700 engineering practitioners and managers across five countries in February 2026. Respondents were split by how often they use AI coding tools. It is vendor-funded research and I am flagging that before quoting it, because the split is the useful part rather than the headline.
Among very frequent AI users, 96% work evenings or weekends several times a month, against 66% of occasional users. 47% report more manual rework after generation, against 28%. Mean incident recovery time runs 7.6 hours against 6.3.
Those four numbers are a better burnout instrument than any adoption chart, because they measure downstream consequence rather than upstream activity. Track them by usage cohort and you will see strain forming a quarter before it shows up in attrition.
One caution on the performance-management side. Do not convert these into individual targets, because the moment after-hours work becomes a metric someone is judged on, it stops being reported honestly. The related question of how performance review criteria should change on AI-assisted teams is genuinely unsettled. The answer is not to grade people on output volume.
Where this argument is weakest
I have argued that pace pressure from AI is a live cause of engineering burnout. Here is the case against my own position, stated properly rather than as a hedge.
None of this establishes cause
Every figure above is correlational or self-reported. The same two years that brought AI adoption also brought sustained layoffs, hiring freezes, flatter budgets and repeated reorganisation across the industry. LeadDev's 2025 edition attributed rising burnout mainly to layoffs, survivor guilt and scope creep, with no mention of AI at all, and 65% of respondents reported expanded responsibilities that year.
An engineer working through a reduced team's workload will feel exhausted whether or not a model is writing the first draft. Some of what this post attributes to pace is more honestly attributed to the headcount correction running alongside it. I cannot separate the two from public data.
Two of the sharpest numbers are vendor-funded
The Harness figures come from a company selling delivery tooling. The Atlassian figures come from a company selling developer tooling. Both published sample size, geography and field dates, which is more than most, and both still have a commercial interest in the finding that engineering work is harder than it looks.
The strongest counter-evidence is DORA's own, and it is not small. AI adoption showed a positive relationship with individual effectiveness, code quality and team performance in the 2025 data. If you already run small batches, a quality internal platform and a clear stance on AI use, the evidence says the tools help you and the burnout risk lives elsewhere. Assume that describes you and you will be wrong more often than right. The possibility is real, and I will not pretend otherwise.
If you are the manager who is also exhausted
The 30-point jump among chief technology officers is the finding I keep returning to. Thomas Johnson, a CTO quoted by LeadDev, described the cause as AI giving teams close to unlimited capacity and creating relentless pressure to write detailed specifications fast enough to feed it.
That is a real and under-discussed shift. When execution capacity rises, the binding constraint moves upstream to whoever decides what should be built, and that person is usually already the least protected member of the org chart.
Two things help, and neither is a wellbeing programme. Cut the number of concurrent initiatives rather than trying to specify all of them faster. And write down, once, which decisions you actually own. A fair share of manager exhaustion in 2026 comes from re-litigating decisions that were never yours to make.
Frequently asked questions
Is AI causing developer burnout?
Not directly, on the current evidence. DORA surveyed nearly 5,000 technology professionals in 2025 and found friction and burnout at similar levels for high and low AI adopters once other variables were controlled. What AI reliably changes is delivery pace and output volume. Burnout follows when expectations rise to match that output while review capacity, headcount and on-call rotas stay fixed.
Why do engineers report saving time with AI but still feel overworked?
Because the saved time does not stay saved. Atlassian found 68% of developers saving 10 or more hours a week with AI, and 90% losing 6 or more hours a week to organisational friction such as finding information and switching contexts. DORA observed separately that time saved during generation is often reallocated to verification. The hours move rather than disappear.
What is a sustainable pace for an AI-assisted engineering team?
There is no published benchmark, and anyone quoting one is guessing. The usable test is behavioural rather than numeric. Ask three questions. Does a meaningful share of your team work evenings or weekends several times a month? Does review sit with two or three people? Do concurrent work items per engineer keep climbing? A yes to any means your pace is above what the team can sustain, whatever the velocity chart says.
How should engineering managers measure burnout risk?
Measure downstream consequences rather than tool adoption. Four indicators work well: frequency of evening and weekend work, manual rework after code generation, concentration of review load across named individuals, and mean incident recovery time. Track them by AI usage cohort and review them quarterly. Never convert them into individual performance targets, because the moment they are graded they stop being reported honestly.
Does using AI coding tools mean we need fewer engineers?
The evidence does not support that conclusion yet. AI raises generation throughput, and DORA found it also raises delivery instability, which increases downstream work in review, testing and incident response. Cutting headcount on the strength of generation speed alone shifts load onto the people who verify and operate the system. That is the exact mechanism producing the exhaustion figures in this post.
Where to start this week
Pick one measurement and one constraint. Do not attempt a programme.
For the measurement, pull the last 8 weeks of merged pull requests and count reviews by reviewer. If more than half sit with fewer than a quarter of the team, you have found your bottleneck and you did not need a survey to find it.
For the constraint, set a concurrent work-in-progress limit per engineer at the next planning session and hold it for one sprint. It will make your queue visibly longer, which is the point. A queue you can see is a resourcing conversation. A queue hidden inside people's evenings is an attrition event with a six-month fuse.
References
- World Health Organization, Burn-out an occupational phenomenon: International Classification of Diseases, 28 May 2019. Used for the ICD-11 definition and its three dimensions.
- Google Cloud, Announcing the 2025 DORA Report, 2025. Used for the 90% adoption figure, respondent count, the amplifier finding and the team archetypes.
- DORA, Balancing AI tensions: moving from AI adoption to effective SDLC use, 2025. Used for verification overhead, reviewer burden and the developer interview quotes.
- Stack Overflow, 2025 Developer Survey, AI section, 48,945 respondents. Used for adoption, trust and the two frustration figures.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Used for the randomised trial result and its stated limitations.
- LeadDev, AI coding is addictive. Engineers are paying the price, 30 June 2026. Used for the Engineering Leadership Report 2026 exhaustion figures and both practitioner quotes.
- Harness, 2026 State of DevOps Modernization Report, conducted by Coleman Parkes, 700 respondents, February 2026. Used for the usage-cohort comparisons.
- Atlassian, State of Developer Experience 2025. Used for hours saved and hours lost to organisational friction.
The weakest part of this source base is that almost every burnout figure is self-reported by a self-selected sample. Two of the sharpest comparisons are vendor-funded. Self-reported exhaustion is not a clinical diagnosis, and the METR result shows that developer self-assessment of their own speed can be wrong by nearly 40 percentage points. Read every number here as directional.
Related reading