From Sanskriti Khandelwal | Product & Market Analysis

Developer Trust in AI Fell to 29% While Usage Hit 84%. Nobody Is Buying on Trust

On this page

29% of developers say they trust the accuracy of AI output. 84% use the tools or plan to. Both figures come from the same Stack Overflow survey of more than 49,000 developers, and both moved last year: trust down from 40%, usage up from 76%. Developer trust in AI is falling while adoption climbs, and the two are not in conflict.

Key takeaways

  • Trust fell 11 points in the same year usage rose 8. Stack Overflow's 2025 survey of more than 49,000 developers puts trust in AI accuracy at 29%, down from 40%, against 84% using or planning to use the tools, up from 76%.
  • The two flagship surveys ask different questions and get opposite readings. Google's DORA study of nearly 5,000 professionals found 30% with little to no trust in AI-generated code. Stack Overflow found 46% distrusting accuracy. Both are defensible.
  • Saved time is being moved, not returned. DORA's own reading is that generation time saved gets re-allocated to verification. METR's randomised trial measured experienced developers running 19% slower with AI while believing they were 20% faster.
  • Vendor revenue is decoupled from user trust because the buyer is not the user. The venture money is already flowing to inspection rather than generation, which is the clearest signal of what the market thinks it needs next.
29%Developers who trust the accuracy of AI output, down from 40% a year earlier. Source: Stack Overflow 2025 Developer Survey.
84%Developers using or planning to use AI tools, up from 76% in 2024. Source: Stack Overflow 2025 Developer Survey.
19%Measured slowdown for experienced developers allowed to use AI on repositories they knew well. Source: METR, July 2025.

The short answer

Developer trust in AI accuracy fell to 29% in Stack Overflow's 2025 survey, down from 40%, while 84% of developers now use or plan to use AI tools, up from 76%. Trust is falling because exposure is rising. Developers hand these tools the work where a mistake is cheap and checkable, and withhold the work where a mistake is expensive and invisible.

What the 29% figure actually measures

Stack Overflow fielded its 2025 Developer Survey across more than 49,000 developers in 177 countries and published the results in December 2025. One question asks how much you trust the accuracy of the output from AI tools. The answer distribution is more interesting than the headline pulled from it.

3.1% say they highly trust it. 29.6% somewhat trust it. On the other side, 26.1% somewhat distrust and 19.6% highly distrust, which is where the widely quoted 46% distrust figure comes from.

The breakdown matters more than the headline

The 29% in most coverage maps to the somewhat-trust band on its own. Add the 3.1% who highly trust and you get 33%. That is a real difference if you are putting the number in a board deck, and almost nobody states which cut they are quoting.

The residual is the part worth staring at. Roughly 22% of respondents fall outside all four bands. On a direct question about whether your daily tools produce correct output, close to 1 developer in 5 declines to take a position at all.

Favourability fell further than trust. It went from 72% in 2024 to 60% in 2025, with the survey's own breakdown giving 59.7% combined favourable. That leaves a 27 point gap between developers who feel positive about AI and developers who trust its accuracy. They like the tools and do not believe them.

Usage went up. Everything else went down. Stack Overflow Developer Survey, 2024 and 2025. Percentage of respondents. 2024 2025 76% 84% Use or plan to use AI 72% 60% Favourable toward AI tools 40% 29% Trust the accuracy of output Two data points per line. This is a direction of travel, not a trend line.
The distance between the top line and the bottom line is the whole story. Usage is not a proxy for belief, and it has not been for at least a year.

The second survey disagrees, and the disagreement is the question

Run a similar question through a different study and you get an almost opposite reading. Google's DORA programme surveyed nearly 5,000 technology professionals for its 2025 State of AI-assisted Software Development report. 90% use AI at work. 30% report little to no trust in AI-generated code.

Read carelessly, DORA says 70% of professionals trust AI code while Stack Overflow says 33% do. That is a 37 point gap between two credible studies published within months of each other.

Two questions, two populations

They are not measuring the same object. DORA asks about trust in AI-generated code on a scale where "somewhat" counts as trust. Stack Overflow asks about the accuracy of AI output, which is a stricter claim and a stricter word.

The populations differ too. Stack Overflow's respondents are people who visit Stack Overflow, a site whose traffic has a direct stake in whether AI answers turn out to be correct. That does not invalidate the finding. It does mean the level should be treated as directional and the movement treated as the signal.

My reading is that the trend is the only part either survey can carry safely. Both show trust flat or falling in a year when usage rose sharply. Agreement across two different questions, two different samples and two different sponsors is worth more than either number on its own.

JetBrains gives the third reading

JetBrains surveyed 24,534 developers across 194 countries for its 2025 State of Developer Ecosystem. 85% use AI regularly. Only 44% report AI as fully or partially integrated into their workflow.

That 41 point gap is the same phenomenon wearing different clothes. Developers open these tools every day and decline to wire them into the parts of the workflow that would run without supervision. Their top stated concern is code quality, at 23%, ahead of privacy and security at 13%.

Where you cut the scale decides the headline Trust in the accuracy of AI output. Stack Overflow 2025 Developer Survey, 49,000+ respondents. 32.7% trust at all Highly trust3.1% Somewhat trust29.6% Neither, computed residual21.6% Somewhat distrust26.1% Highly distrust19.6% The quoted 29% is the second band alone. The quoted 46% is the bottom two combined. The residual is arithmetic on published percentages, not a figure Stack Overflow reports directly.
Notice how few people are in the top band. 3.1% highly trusting is the number a vendor should be trying to move, and it is the one nobody quotes.
Three 2025 developer surveys, and what each one actually asked
StudySampleUsage findingTrust finding, and the question behind it
Stack Overflow 2025 Developer Survey49,000+ developers, 177 countries84% use or plan to use AI tools29% trust the accuracy of output. Asks about accuracy specifically.
DORA, State of AI-assisted Software Development 2025Nearly 5,000 technology professionals90% use AI at work30% have little to no trust in AI-generated code. Broader scale, softer word.
JetBrains State of Developer Ecosystem 202524,534 developers, 194 countries85% use AI regularlyDoes not ask trust directly. 44% report AI integrated into the workflow.

All three recruit from their own audiences, so none is a probability sample of working developers. Compare the direction of travel across them, not the absolute levels.

Adoption is not endorsement, and the agent numbers prove it

The agent figures are the cleanest evidence that usage and belief are separate variables. 14.1% of Stack Overflow respondents use AI agents daily. 37.9% have no plans to use them at all.

Among developers who do use agents, 87% agree that accuracy is a challenge and 81.4% flag security and privacy. Those are not the numbers of a category that has won an argument. They are the numbers of a category being trialled under close supervision.

The same survey found 77% of developers do not use vibe coding professionally, even as roughly half said agents had affected their work. Exposure is near universal. Delegation is not.

Where developers refuse to hand over control

The survey also asks which parts of the workflow developers will not use AI for. This refusal map is the most useful chart in the whole release and it is almost never quoted.

75.8% do not plan to use AI for deployment and monitoring. 69.2% refuse for project planning. 58.7% refuse for committing and reviewing code. Resistance is lowest for searching for answers, at 19.6%, and for writing code, at 28.9%.

That pattern is coherent and it is not fear. Developers hand AI the work where a mistake is cheap and immediately visible. They keep the work where a mistake is expensive and surfaces later. Read as a risk allocation rather than a trust deficit, it is exactly what you would want a professional to do.

The refusal map: where developers will not let AI in Share of respondents saying they do not plan to use AI for each task. Stack Overflow 2025. Deployment and monitoring75.8% Project planning69.2% Committing and reviewing code58.7% Writing code28.9% Searching for answers19.6% Blue is where a mistake is cheap and visible within seconds. Red is where a mistake is expensive and shows up days later, in production.
Every bar above 50% is a workflow stage where the failure is delayed. That, and not sentiment, is what predicts refusal.

The verification tax is where the trust went

Falling trust is usually reported as a mood. It is closer to a cost line, and the survey data names the cost precisely.

Almost right is the expensive failure mode

66% of Stack Overflow respondents name AI solutions that are almost right, but not quite, as a frustration. 45.2% say debugging AI-generated code is more time-consuming than it is worth. 20% report feeling less confident in their own problem solving, and 16.3% say the logic is hard to follow.

A wrong answer is cheap. It fails fast, you throw it away, you move on. A nearly right answer is expensive, because it survives the first read, enters the codebase, and fails somewhere further downstream where the context has gone.

Only 4.4% of respondents think AI handles complex tasks very well, while 39.6% rate that performance poorly or very poorly. Trust and usage can move in opposite directions because the tool is genuinely useful and its failure mode is genuinely costly at the same time.

DORA's own framing of the 2025 data says the same thing in fewer words. Time saved during initial generation is often re-allocated to verification overhead. The hour does not come back. It changes department.

METR measured the size of it once

METR ran a randomised controlled trial with 16 experienced open-source developers across 246 issues. The repositories were ones they already knew, averaging more than 22,000 GitHub stars and a million lines of code.

Developers using AI took 19% longer to complete their issues. Before the trial they forecast a 24% speedup. After finishing slower, they still estimated AI had made them 20% faster.

METR is careful about what this does not show, and so am I. 16 developers on mature repositories is not a claim about the profession, and the authors say so directly. The figure I would generalise is not the 19%. It is the perception gap of roughly 39 points, which survived direct personal experience of the opposite result.

That gap is the reason self-reported productivity surveys should not settle this question, and it is a smaller version of the measurement problem covered in the piece on why AI productivity gains keep failing to show up in the aggregate data.

What the codebase shows when nobody is being asked

Surveys record opinions. Repository telemetry records behaviour, and it is moving in the same direction.

GitClear and GitKraken analysed 623 million real-world code changes from 2023 to 2026. Duplicated code blocks rose 81%. Copy and paste within a single commit rose 41%. Moved lines, the standard signal for refactoring, fell 70%. Legacy refactoring fell 74% since 2023. Cross-file function calls, a proxy for reuse, fell 35%.

Read that as correlation, because that is what it is. The window covers a hiring contraction and a rate cycle as well as an AI adoption wave, and the dataset is repositories that use GitClear rather than a random sample of software. Any one of those metrics could be explained away.

The part that is harder to dismiss is the consistency. Five independent maintainability signals moved the same way over four years. I would not publish this as proof that AI degrades code. I would publish it as evidence that whatever is happening to maintainability is not being caught by the review process most teams still run.

The money is already moving to the verification layer

If falling trust were only a sentiment problem, the market would be spending on reassurance. It is spending on inspection.

In September 2025, CodeRabbit closed a $60 million Series B after growing revenue 10 times in a year. Graphite had raised $52 million in March 2025 and was acquired by Cursor in December 2025. Greptile raised a $25 million Series A led by Benchmark, taking its total to $30 million at a reported $180 million valuation.

The stated reason is mechanical rather than emotional. Code volume has risen faster than review capacity, and the systems built to validate code before deployment cannot keep pace with agents that write continuously. Individual distrust shows up at company level as a procurement line item.

That is the shape of the opportunity for anyone building here. The generation layer is crowded and commoditising. The layer that tells you which generated output you can safely skip reading is not, and it is the closer analogue to the workflow positions described in the analysis of why workflow depth outlasts data advantage.

The buyer is not the user

This is why vendor revenue and user trust can diverge for years without anything breaking. The developer forms the trust judgement. The engineering leader signs the contract, on a throughput argument, at a seat price low enough that the decision does not have to be right.

A $20 monthly seat does not require belief. It requires the absence of a strong objection. That is a far lower bar, and it is the bar most AI developer tools currently clear, which is a version of the dynamic examined in the breakdown of what happens to seat-based pricing when the seat stops being the unit of value.

The exposure arrives at renewal, when a finance team asks what changed. At that point the vendor needs a number, and self-reported time savings will not survive contact with a CFO who has read how thin the measured returns on generative AI have actually been.

What to do about it, on either side of the table

The trust decline is not a brand problem to be handled with messaging. It is a product specification, and it is fairly specific about what it wants.

The claim most vendors make, against the claim the data supports
The pitch todayThe pitch this data supportsWhy the first one still works
"Writes 40% of your code""40% of what it wrote is still in the repository, untouched, 30 days later"Generation volume is easy to instrument and easy to demo in a sales call.
"30% suggestion acceptance rate"Rework rate on AI-authored changes, against your human-authored baselineAcceptance is already tracked in the editor. Survival needs repository history.
"Saves 6 hours per developer per week"Where the saved hours went, review and debugging includedSelf-reported time savings survey extremely well, as METR demonstrated.
"Trusted by 60 Fortune 500 companies"The failure classes the tool gets wrong, named in the documentationLogos clear procurement. Publishing your own failure modes is uncomfortable and slower.

The third column is not a throwaway. The left-hand claims are winning right now, which is why the right-hand ones are a bet on where the market goes rather than a description of where it is.

If you sell one of these tools, publish one number with a stated sample, a stated window and a stated exclusion list. Acceptance rate is not that number, because acceptance is measured at the moment of least information, roughly two seconds after the suggestion appears. Survival rate is closer: what share of AI-authored lines is still present and unedited a month later.

If you buy one, the metric is rework rather than adoption. Pull your merged changes, mark which were AI-authored, and measure how many were modified again within 14 days. Compare that to the same figure for human-authored changes over the same period.

Most teams cannot run that comparison because they never recorded the baseline, which is the same gap that makes the build-or-buy question so hard to settle honestly. That decision is worked through in the piece on when building your own coding agent actually pays.

Where this argument is weakest

Four problems with everything above, stated plainly rather than buried.

Every survey quoted here is self-selected. Stack Overflow, DORA and JetBrains all recruit from their own audiences, and each audience has a relationship with AI tooling that a random sample of developers would not share. The levels are unreliable even where the directions agree.

Falling trust may be better calibration

The pessimistic reading is that output quality is not improving fast enough to keep pace with exposure. The optimistic reading is that developers are getting more accurate about where these tools fail, after a year of using them daily.

Those two look identical in a survey response. If the second is right, falling trust is a healthy signal, and the tools are being used more effectively by people who have stopped expecting them to be correct by default.

I lean toward the second reading, with one caveat. The refusal map shows developers withholding AI from the highest-value stages of the workflow, and improved calibration on its own does not explain that pattern.

The reporting is inconsistent with itself

Stack Overflow's own results page and its own summary blog report two frustration figures with the numbers attached to different statements. 66% and 45% swap places depending on which page you read. I have used the results page throughout, because it sits closer to the raw data, and you should know that choice was made.

The 22% residual I calculated for the trust question is my arithmetic on published percentages, not a figure Stack Overflow states. The DORA report's full trust breakdown sits inside a PDF I could not open in full, so the 30% figure comes from DORA's own summary page rather than the underlying table.

Finally, METR has since revised its productivity experiment design. That is what a serious research group does when it learns something, and it is also a reason not to build a thesis on a single trial of 16 developers.

Frequently asked questions

How many developers trust AI coding tools in 2025?

29% of developers trust the accuracy of AI output, according to Stack Overflow's 2025 Developer Survey of more than 49,000 respondents. That is down from 40% the previous year. Combined distrust reached 46%, and only 3.1% said they highly trust the output. Google's DORA study of nearly 5,000 professionals reported a softer figure, with 30% expressing little or no trust in AI-generated code.

Why is developer trust in AI falling while usage rises?

Because exposure and confidence are different things. More developers now use these tools daily, which means more of them have met the failure mode directly. The most cited frustration is output that is almost right but not quite, named by 66% of respondents. Usage keeps rising because the tools remain useful for cheap, checkable work even when nobody believes them by default.

What did the Stack Overflow 2025 Developer Survey say about AI?

It surveyed more than 49,000 developers in 177 countries. 84% use or plan to use AI tools, up from 76%. Trust in accuracy fell to 29% from 40%, favourability fell to 60% from 72%, and 46% actively distrust the output. 14.1% use AI agents daily while 37.9% have no plans to use agents at all.

Does AI actually make developers faster?

The evidence is mixed and thinner than the marketing suggests. METR's randomised trial found 16 experienced developers took 19% longer on their own repositories with AI available, while estimating afterwards that they had been 20% faster. DORA's 2025 data associates AI adoption with higher throughput but also higher delivery instability. Neither result generalises cleanly to every team.

What is the biggest frustration with AI coding tools?

Output that is almost right, cited by 66% of Stack Overflow respondents. A wrong answer is cheap because it fails immediately. A nearly right answer is expensive because it passes the first read and fails later, in a context where the original reasoning has been lost. 45.2% separately said debugging AI-generated code is more time-consuming than it is worth.

What should AI coding tool vendors do about falling trust?

Treat it as a product specification rather than a messaging problem. Replace generation-volume claims with survival metrics: what share of the code the tool wrote is still present and unedited 30 days later, with the sample and window stated. Venture funding is already moving toward code review and validation, which suggests the market has priced verification as the next scarce thing.

Where to start this week

Two things, and both fit inside an afternoon.

If you sell a developer AI tool, take one claim off your homepage and replace it with a survival rate. What share of the code your product wrote is still in the repository, unedited, 30 days later, with the sample size and window published alongside it. If you cannot compute that figure today, you have just found the next thing to build.

If you buy one, pull the last 90 days of merged changes and mark which were AI-authored. Measure how many were touched again within 14 days, then run the same measurement on human-authored changes. That one ratio will settle more internal arguments than every survey quoted in this post.

If you are deciding what to build next

The verification layer argument connects directly to two other pieces: when building your own coding agent actually pays, and why workflow depth outlasts a data advantage.

References

  1. Stack Overflow, 2025 Developer Survey, AI section, December 2025. Used for all adoption, trust, favourability, frustration, agent and refusal figures.
  2. Stack Overflow, Developers remain willing but reluctant to use AI, 29 December 2025. Used for the year-on-year trust and favourability comparison.
  3. DORA, Balancing AI tensions: moving from AI adoption to effective SDLC use, 2026. Used for the 30% low-trust figure and the verification overhead framing.
  4. Google Cloud, Announcing the 2025 DORA Report, 2025. Used for adoption rate, sample size and the throughput and instability relationship.
  5. METR, Measuring the impact of early-2025 AI on experienced open-source developer productivity, 10 July 2025. Used for the 19% slowdown and the perception gap.
  6. JetBrains, The State of Developer Ecosystem 2025, October 2025. Used for the 85% usage and 44% integration figures.
  7. LeadDev, Code maintainability plummets in the AI coding era, 2026. Used for the GitClear and GitKraken maintainability metrics.
  8. SiliconANGLE, Greptile bags $25M in funding to take on CodeRabbit and Graphite in AI code validation, 23 September 2025. Used for all code review funding figures.

The weakest part of this source base is that every survey cited recruits from its sponsor's own audience, so none is a probability sample of working developers. The GitClear figures reach this post through secondary coverage because the original research page returned an access error, and they should be upgraded to the primary report when it becomes reachable.

SK
Sanskriti Khandelwal
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading