From Sanskriti Khandelwal | Product & Market Analysis
Agent Payback by Function: Six Studies Say Under 6 Months, One Says 15
On this page
Seven published studies put a payback number on enterprise AI agents. Six of them say under six months, and between them they cover six different business functions. The seventh, built on 420 survey responses rather than four customer interviews, says 15 months. The pattern in agent payback data is not about function. It is about who paid for the study.
Key takeaways
- Six commissioned studies across six functions all land on the same payback figure. Customer service, contact centre, engineering, finance, IT operations and marketing. Reported ROI ranges from 207% to 449%. The payback line reads "under six months" in every one.
- The study with the largest sample reports the longest payback. Forrester's January 2026 analysis of Microsoft's agentic AI solutions modelled $44.5 million of benefits against $20.2 million of costs, giving 120% ROI and a 15-month payback.
- Buyers describe a different world. Forrester surveyed over 1,400 AI decision-makers: 13% reported a positive EBITDA impact and fewer than a third could link AI to P&L. Nearly half still expect payback inside a year.
- The by-function payback table circulating online has no primary source. Every restatement of "4.1 months for customer service" traces back to another blog post. Bain's own 2026 survey of 951 companies publishes no payback-by-function figures at all.
The short answer
There is no reliable median payback period for AI agents by function. The studies that publish one are vendor-commissioned models of a designed composite company, and they cluster under six months whatever the function. Independent buyer surveys report far weaker returns and give no function breakdown at all.
That is the finding. The rest of this post is the evidence behind it, and what to put in a business case instead.
What the seven published studies actually report
If you want a payback figure with a written methodology attached, one genre dominates. It is the Forrester Total Economic Impact study, and every one is paid for by the vendor whose product it measures.
That is not a secret. Each study says so on its cover. It does mean the public evidence base for agent payback is sponsored research, which is worth holding in mind before any of these numbers reaches a slide.
| Function | Study | ROI | Payback | Evidence base |
|---|---|---|---|---|
| Customer service | Salesforce Agentforce, November 2025 | 396% | Under 6 months | 6 interviews. Composite runs 500,000 cases a year. |
| Contact centre | Google Cloud Customer Engagement Suite, May 2025 | 207% | Under 6 months | 4 interviews. Composite is a $10bn firm, 6,000 agents. |
| Engineering | GitLab Duo Agent Platform, July 2026 | 400% | Under 6 months | 4 interviews. Composite is a $3bn firm, 150 users. |
| Finance and accounting | FloQast, 2026 | 275% | Under 6 months | Composite. $2.8m saved, close time halved. |
| IT operations | LogicMonitor Edwin AI, 2026 | 313% | 6 months or less | Composite. Alert triage and incident resolution. |
| Marketing | Insider One, 2026 | 449% | Under 6 months | Composite global retailer. $13.9m NPV. |
| Cross-function platform | Microsoft agentic AI solutions, January 2026 | 120% | 15 months | 8 interviews across 6 firms, plus 420 survey respondents. |
All seven are Forrester Consulting studies commissioned by the vendor named. ROI figures are three-year, risk-adjusted and in present value terms. Where payback is published as a band, the band is reproduced.
How a Total Economic Impact study is built
Forrester interviews a small number of customers the vendor introduces. It then builds a composite organisation that blends those customers into one modelled company, applies a risk adjustment to the benefits, and discounts three years of cash flows to present value.
The method is transparent and the assumptions are printed. It is still a model of an organisation that does not exist, populated by customers who agreed to speak about a product they chose to keep buying.
What "payback" means inside these models
Payback here is the month in which the composite's cumulative risk-adjusted benefits exceed its cumulative costs. It is not the month your finance team stops seeing a net outflow.
The difference matters. The composite's costs are licence fees, professional services and some internal labour. Your costs include a data integration project nobody scoped, and a quarter of people quietly working around the thing.
Why six different functions produce the same number
A payback figure has three inputs: the benefit, the cost, and the shape of the curve between them. In a commissioned model, all three are chosen rather than observed.
The composite is designed, not sampled
Look at the evidence base column above. Four interviews. Six interviews. Four again. These are not samples in any statistical sense, and Forrester does not claim they are.
A composite built from four references will reflect four deployments that went well enough for the customer to speak on record. Deployments that stalled are not in the room. That selection alone is enough to compress payback.
The 15-month outlier has the largest sample
The Microsoft study is interesting precisely because it breaks the pattern. It combines 8 interviews across 6 organisations with a survey of 420 people who had used the products, then models a composite with $2.5 billion of revenue and 10,000 employees.
Its answer is $44.5 million of three-year benefits against $20.2 million of costs, a net present value of $24.2 million, and 15 months to payback. Wider evidence, broader scope, longer payback. That relationship is the most useful signal in the table.
Nobody commissions a study that lands at 14 months
Here is the part worth saying plainly. A vendor pays for one of these studies to support a sales motion, and the sales motion needs the number to be short.
None of this requires anyone to lie. It requires only that studies which come back weak never get published, while studies that come back strong get a landing page and a press release. Publication bias needs no bad actors, and it explains a tight cluster better than any claim about functions does.
What buyers report when nobody is paying for the answer
Set the commissioned studies aside and look at survey work sold to buyers rather than sponsored by sellers. The picture changes shape completely.
Forrester surveyed over 1,400 global AI decision-makers for its January 2026 analysis. Only 13% reported a positive EBITDA impact, and fewer than a third could link AI contributions to P&L. In the same survey, nearly half expect payback within a year while only 14% commit to a three-year horizon.
Those two findings sit badly together. Buyers demand a payback window that almost nobody in the same sample has demonstrated. The expectation is a procurement policy, not a measurement, and the commissioned studies are calibrated to clear it.
Bain's Automation and AI Pathfinder Survey 2026, published on 1 June 2026 across 951 global companies, adds the operational version. Nearly 40% of companies that measured AI cost savings came in below 10%, against a target of 11% to 20%. Only 7% run fully autonomous agents in production, and 41% name data access and integration as the top barrier.
Notice what Bain does not report. There is no payback-by-function table in it, and no median payback month for agents.
The by-function table in circulation, and what sits under it
Search for agent payback by department and you meet a specific set of figures within two clicks. A 5.1-month median across functions. 4.1 months for customer service, 6.7 for marketing operations, 9.3 for engineering, 3.4 for sales development, 8.9 for finance and operations. They are usually credited to a Bain agentic AI benchmark, or to BCG and Forrester surveys.
I went looking for the primary source. Every result was another blog post, a statistics roundup or a vendor content page. None of the figures appears on bain.com, and Bain's published 2026 survey work reports different metrics entirely.
What the restatements disagree about
The tell is not that the number repeats. It is that the restatements contradict each other about what the number measures.
One page presents 5.1 months as the cross-function median. Another attributes the same 5.1 months to sales development and marketing automation agents specifically. A third reports a Forrester study of 287 deployments with a median payback of 7.3 months. Those cannot all be true, and six restatements of a figure are still one source at best.
My position is simple: none of those numbers belongs in a board deck. This is the same failure pattern as the widely quoted agent pilot failure rate, which traces to aggregators rather than to any survey. A precise decimal is a rhetorical device as often as it is a measurement.
Engineering, the one function with an independent measurement
There is exactly one function where somebody ran a controlled experiment and published the result without a vendor paying for it.
METR recruited 16 experienced open-source developers and randomised 246 real tasks in repositories they already maintained, using Cursor Pro with Claude 3.5 and 3.7 Sonnet. Before starting, the developers expected AI to speed them up by 24%. They took 19% longer with the tools allowed, and afterwards still believed AI had made them 20% faster.
Set that against the GitLab study above, which reports a 20% individual developer productivity gain worth $7.4 million over three years, plus 80% faster onboarding and a code migration compressed from eight months into two.
Two studies asking two different questions
Both results can hold at once, and reading them as a contradiction is the mistake.
METR measured expert maintainers doing familiar work in codebases they know intimately. That is where an agent has least context advantage and most review overhead. GitLab's benefits sit in onboarding, migration and security remediation, which are the unfamiliar, high-volume tasks where the trade runs the other way.
So stop asking whether agents help engineering. Ask which engineering work, because the sign of the effect flips inside one function. That distinction also separates a real tooling decision from a licence purchase, as the build versus buy comparison for coding agents sets out.
Where this argument is weakest
Three places, and the first would change the most if I am wrong.
Commissioned is not the same as wrong
Total Economic Impact studies publish their methodology, name their composite, state their discount rate and apply an explicit risk adjustment to benefits. That is more disclosure than almost any free industry survey offers.
The bias sits in which customers get interviewed and which studies reach publication, not in the arithmetic. A reader who discounts them entirely throws away the most detailed cost models available in public.
One trial is not a verdict on a function
METR is 16 developers, one tooling generation, and open-source repositories with unusually high review standards. METR says so itself, and warns against reading the result as evidence that AI fails developers generally.
Agent capability also moved between early 2025 and now. A 2026 rerun could land on the other side of zero, and I would not be surprised if it did.
The third weakness is mine. I could not find a primary source for the circulating by-function table, and that is not the same as proving none exists. A paywalled benchmark could carry those figures. What I can say is narrower: no public page I opened contains them with a methodology attached, and the pages repeating them disagree about what they measure.
How to build a payback number your CFO can check
The useful move is not to find a better benchmark. It is to build a small number you own, which beats a large number you borrowed.
Start with a unit of work that already carries a price in your business. Cost per support contact. Cost per invoice processed. Fully loaded cost per qualified meeting booked. If a function has no priced unit of work, it has no payback number either, and that is worth discovering before procurement rather than after.
| Cost line | In the commissioned model | What to add before you sign |
|---|---|---|
| Licence and platform fees | Counted, usually accurately | Model the usage tier you land in at month 12, not month 1. |
| Implementation and integration | Counted as professional services | Internal engineering time on data access, the barrier 41% of Bain respondents named first. |
| Change management and training | Sometimes counted, often thin | Hours per affected employee, times the number who actually change how they work. |
| Review and rework | Rarely counted | Time spent checking agent output, which is where the METR slowdown came from. |
| Incident and error handling | Not counted | A budgeted allowance for the failures that reach a customer. |
| The counterfactual | Never counted | What the same budget and the same six months would return elsewhere. |
The last row is the one finance teams ask about and the one no vendor study can answer, because a comparison against the next best use of the money is not a product benefit.
Record the baseline before the pilot starts, with a date on it. A baseline reconstructed afterwards gets reconstructed to fit, and the METR result shows how confidently people misremember their own throughput. The same discipline applies to the contract, where the clauses a CFO should require before signing an AI deal decide whether a disappointing month is renegotiable or simply invoiced.
Then price the thing properly. Agent pricing is usually consumption-shaped, so a model built on a flat licence assumption will be wrong by month nine. The conversation-by-conversation arithmetic behind Agentforce pricing shows how fast a per-unit rate compounds at real volumes. It is the small version of the question the hyperscalers face in the payback math behind 2026 AI capital spending, and the small version is the one you can answer.
Frequently asked questions
What is the average payback period for AI agents?
No trustworthy average exists. Six published Forrester Total Economic Impact studies report payback in under six months, but each is commissioned by the vendor and models a composite organisation built from four to six customer interviews. The one study using a 420-person survey reports 15 months. Independent buyer surveys do not publish a payback average at all, and only 13% of decision-makers report a positive EBITDA impact.
Which business function gets the fastest ROI from AI agents?
The published evidence cannot answer this. Studies covering customer service, contact centre, engineering, finance, IT operations and marketing all report payback in the same under-six-months band, so the data has no power to rank functions. Circulating tables that rank sales fastest and finance slowest do not trace to a primary source. Judge by whether the function has a priced unit of work instead.
Are Forrester Total Economic Impact studies reliable?
They are transparent and sponsored at the same time. Each publishes its methodology, composite organisation, discount rate and risk adjustments, which is more disclosure than most free surveys offer. The bias is in selection: the interviewed customers are supplied by the vendor, and weak results are unlikely to reach publication. Treat them as detailed cost models rather than as measurements of a market.
Why do AI agent ROI figures vary so much?
Because the inputs are chosen rather than observed. Reported ROI in the seven studies ranges from 120% to 449%, driven by the size of the modelled composite, which benefits are counted and how aggressively they are risk-adjusted. Scope matters most: single-function point solutions report the highest returns, while the one platform-wide study reports the lowest at 120%.
How do I calculate payback for an AI agent project?
Pick a unit of work that already has a price, such as cost per support contact or per invoice processed. Record the baseline before deployment, with a date. Add licence fees at the usage tier you expect by month 12, internal integration engineering, training hours, review and rework time, and an allowance for incidents. Then compare the result against the next best use of the same money.
Where to start this week
Take the agent business case currently sitting in your pipeline and do two things to it.
First, find every payback figure in it and write the primary source next to each one, with the sample size and who paid for the study. Anything you cannot trace to a document you opened yourself comes out of the deck. On current evidence that removes most of the numbers, which is the point.
Second, write down the unit of work the agent is supposed to make cheaper, and today's price for one of them. If nobody in the room knows that price, you do not have a business case yet. You have a procurement request with a chart on it.
Related analysis
For the wider version of this question, see where measurable AI return has actually shown up, and the six failure modes that stop agent pilots reaching production.
References
- Forrester, Three Questions That Will Define AI In 2026, 26 January 2026. Survey of over 1,400 global AI decision-makers. Used for the EBITDA, P&L and payback horizon figures.
- Bain & Company, Your AI Budget Is Growing. Your Returns Aren't., 1 June 2026. Automation and AI Pathfinder Survey 2026, 951 companies. Used for cost savings, autonomy and data barriers.
- METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Used for the 19% slowdown and the expectation gap.
- Forrester Consulting, The Total Economic Impact of Microsoft's Agentic AI Solutions, January 2026. Used for the 120% ROI and 15-month payback.
- Forrester Consulting, The Total Economic Impact of Salesforce Agentforce, November 2025. Used for the customer service row.
- Forrester Consulting, The Total Economic Impact of Google Cloud Customer Engagement Suite, May 2025. Used for the contact centre row.
- GitLab, Study finds organizations can achieve 400% ROI with GitLab Duo Agent Platform, 16 July 2026. Used for the engineering row.
- Vendor summaries of commissioned Forrester studies for LogicMonitor Edwin AI, FloQast and Insider One, 2026. Used for the IT operations, finance and marketing rows.
Weakest thing about this source base: every payback figure in the table comes from research a vendor paid for, and three rows were read from vendor summaries rather than the full study documents. The two independent sources here publish no payback-by-function data at all.
Related reading