From Sanskriti Khandelwal | Product & Market Analysis

Agent Payback by Function: Six Studies Say Under 6 Months, One Says 15

On this page

Seven published studies put a payback number on enterprise AI agents. Six of them say under six months, and between them they cover six different business functions. The seventh, built on 420 survey responses rather than four customer interviews, says 15 months. The pattern in agent payback data is not about function. It is about who paid for the study.

Key takeaways

  • Six commissioned studies across six functions all land on the same payback figure. Customer service, contact centre, engineering, finance, IT operations and marketing. Reported ROI ranges from 207% to 449%. The payback line reads "under six months" in every one.
  • The study with the largest sample reports the longest payback. Forrester's January 2026 analysis of Microsoft's agentic AI solutions modelled $44.5 million of benefits against $20.2 million of costs, giving 120% ROI and a 15-month payback.
  • Buyers describe a different world. Forrester surveyed over 1,400 AI decision-makers: 13% reported a positive EBITDA impact and fewer than a third could link AI to P&L. Nearly half still expect payback inside a year.
  • The by-function payback table circulating online has no primary source. Every restatement of "4.1 months for customer service" traces back to another blog post. Bain's own 2026 survey of 951 companies publishes no payback-by-function figures at all.
6 of 7Published agent ROI studies reporting payback under 6 months, across six functions. Source: Forrester Total Economic Impact studies commissioned by Salesforce, Google Cloud, GitLab, FloQast, LogicMonitor and Insider, 2025 to 2026.
13%Share of AI decision-makers reporting a positive EBITDA impact from AI. Source: Forrester, January 2026.
19%How much longer experienced developers took on real tasks with AI tools allowed. Source: METR, July 2025.

The short answer

There is no reliable median payback period for AI agents by function. The studies that publish one are vendor-commissioned models of a designed composite company, and they cluster under six months whatever the function. Independent buyer surveys report far weaker returns and give no function breakdown at all.

That is the finding. The rest of this post is the evidence behind it, and what to put in a business case instead.

What the seven published studies actually report

If you want a payback figure with a written methodology attached, one genre dominates. It is the Forrester Total Economic Impact study, and every one is paid for by the vendor whose product it measures.

That is not a secret. Each study says so on its cover. It does mean the public evidence base for agent payback is sponsored research, which is worth holding in mind before any of these numbers reaches a slide.

Published AI agent ROI studies, by function, 2025 to 2026
FunctionStudyROIPaybackEvidence base
Customer serviceSalesforce Agentforce, November 2025396%Under 6 months6 interviews. Composite runs 500,000 cases a year.
Contact centreGoogle Cloud Customer Engagement Suite, May 2025207%Under 6 months4 interviews. Composite is a $10bn firm, 6,000 agents.
EngineeringGitLab Duo Agent Platform, July 2026400%Under 6 months4 interviews. Composite is a $3bn firm, 150 users.
Finance and accountingFloQast, 2026275%Under 6 monthsComposite. $2.8m saved, close time halved.
IT operationsLogicMonitor Edwin AI, 2026313%6 months or lessComposite. Alert triage and incident resolution.
MarketingInsider One, 2026449%Under 6 monthsComposite global retailer. $13.9m NPV.
Cross-function platformMicrosoft agentic AI solutions, January 2026120%15 months8 interviews across 6 firms, plus 420 survey respondents.

All seven are Forrester Consulting studies commissioned by the vendor named. ROI figures are three-year, risk-adjusted and in present value terms. Where payback is published as a band, the band is reproduced.

How a Total Economic Impact study is built

Forrester interviews a small number of customers the vendor introduces. It then builds a composite organisation that blends those customers into one modelled company, applies a risk adjustment to the benefits, and discounts three years of cash flows to present value.

The method is transparent and the assumptions are printed. It is still a model of an organisation that does not exist, populated by customers who agreed to speak about a product they chose to keep buying.

What "payback" means inside these models

Payback here is the month in which the composite's cumulative risk-adjusted benefits exceed its cumulative costs. It is not the month your finance team stops seeing a net outflow.

The difference matters. The composite's costs are licence fees, professional services and some internal labour. Your costs include a data integration project nobody scoped, and a quarter of people quietly working around the thing.

Seven commissioned studies. One payback answer. Reported three-year ROI, with the payback figure each study publishes Marketing449%Under 6 mo Engineering400%Under 6 mo Customer service396%Under 6 mo IT operations313%Under 6 mo Finance275%Under 6 mo Contact centre207%Under 6 mo Platform, all fu…120%15 months The ROI column varies by a factor of nearly four. The payback column barely moves. Source: Forrester Consulting Total Economic Impact studies, 2025 to 2026, each commissioned by the vendor named.
Read the right-hand column, not the bars. Six functions, six sponsors, one answer, and a single outlier in red.

Why six different functions produce the same number

A payback figure has three inputs: the benefit, the cost, and the shape of the curve between them. In a commissioned model, all three are chosen rather than observed.

The composite is designed, not sampled

Look at the evidence base column above. Four interviews. Six interviews. Four again. These are not samples in any statistical sense, and Forrester does not claim they are.

A composite built from four references will reflect four deployments that went well enough for the customer to speak on record. Deployments that stalled are not in the room. That selection alone is enough to compress payback.

The 15-month outlier has the largest sample

The Microsoft study is interesting precisely because it breaks the pattern. It combines 8 interviews across 6 organisations with a survey of 420 people who had used the products, then models a composite with $2.5 billion of revenue and 10,000 employees.

Its answer is $44.5 million of three-year benefits against $20.2 million of costs, a net present value of $24.2 million, and 15 months to payback. Wider evidence, broader scope, longer payback. That relationship is the most useful signal in the table.

Nobody commissions a study that lands at 14 months

Here is the part worth saying plainly. A vendor pays for one of these studies to support a sales motion, and the sales motion needs the number to be short.

None of this requires anyone to lie. It requires only that studies which come back weak never get published, while studies that come back strong get a landing page and a press release. Publication bias needs no bad actors, and it explains a tight cluster better than any claim about functions does.

What buyers report when nobody is paying for the answer

Set the commissioned studies aside and look at survey work sold to buyers rather than sponsored by sellers. The picture changes shape completely.

Forrester surveyed over 1,400 global AI decision-makers for its January 2026 analysis. Only 13% reported a positive EBITDA impact, and fewer than a third could link AI contributions to P&L. In the same survey, nearly half expect payback within a year while only 14% commit to a three-year horizon.

Those two findings sit badly together. Buyers demand a payback window that almost nobody in the same sample has demonstrated. The expectation is a procurement policy, not a measurement, and the commissioned studies are calibrated to clear it.

Bain's Automation and AI Pathfinder Survey 2026, published on 1 June 2026 across 951 global companies, adds the operational version. Nearly 40% of companies that measured AI cost savings came in below 10%, against a target of 11% to 20%. Only 7% run fully autonomous agents in production, and 41% name data access and integration as the top barrier.

Notice what Bain does not report. There is no payback-by-function table in it, and no median payback month for agents.

What buyers require, and what buyers can show Forrester survey of over 1,400 global AI decision-makers, January 2026 Expect payback within 12 months Nearly half Commit to a 3-year payback horizon 14% Report a positive EBITDA impact 13% Bars are shares of the same respondent base. The scale runs from 0 to 100%. Fewer than a third of the same group can link any AI contribution to P&L.
The top bar is a purchasing rule. The bottom bar is an outcome. Most business cases are written as if they were the same measurement.

The by-function table in circulation, and what sits under it

Search for agent payback by department and you meet a specific set of figures within two clicks. A 5.1-month median across functions. 4.1 months for customer service, 6.7 for marketing operations, 9.3 for engineering, 3.4 for sales development, 8.9 for finance and operations. They are usually credited to a Bain agentic AI benchmark, or to BCG and Forrester surveys.

I went looking for the primary source. Every result was another blog post, a statistics roundup or a vendor content page. None of the figures appears on bain.com, and Bain's published 2026 survey work reports different metrics entirely.

What the restatements disagree about

The tell is not that the number repeats. It is that the restatements contradict each other about what the number measures.

One page presents 5.1 months as the cross-function median. Another attributes the same 5.1 months to sales development and marketing automation agents specifically. A third reports a Forrester study of 287 deployments with a median payback of 7.3 months. Those cannot all be true, and six restatements of a figure are still one source at best.

My position is simple: none of those numbers belongs in a board deck. This is the same failure pattern as the widely quoted agent pilot failure rate, which traces to aggregators rather than to any survey. A precise decimal is a rhetorical device as often as it is a measurement.

Engineering, the one function with an independent measurement

There is exactly one function where somebody ran a controlled experiment and published the result without a vendor paying for it.

METR recruited 16 experienced open-source developers and randomised 246 real tasks in repositories they already maintained, using Cursor Pro with Claude 3.5 and 3.7 Sonnet. Before starting, the developers expected AI to speed them up by 24%. They took 19% longer with the tools allowed, and afterwards still believed AI had made them 20% faster.

Set that against the GitLab study above, which reports a 20% individual developer productivity gain worth $7.4 million over three years, plus 80% faster onboarding and a code migration compressed from eight months into two.

Two studies asking two different questions

Both results can hold at once, and reading them as a contradiction is the mistake.

METR measured expert maintainers doing familiar work in codebases they know intimately. That is where an agent has least context advantage and most review overhead. GitLab's benefits sit in onboarding, migration and security remediation, which are the unfamiliar, high-volume tasks where the trade runs the other way.

So stop asking whether agents help engineering. Ask which engineering work, because the sign of the effect flips inside one function. That distinction also separates a real tooling decision from a licence purchase, as the build versus buy comparison for coding agents sets out.

The measurement and the memory point in opposite directions METR randomised trial, 16 experienced developers, 246 real tasks, 2025 +24% Expected before the trial -19% Measured +20% Believed after the trial faster slower Self-reported productivity gains are not evidence. The same people who slowed down reported speeding up.
This is why a post-deployment survey is the weakest possible input to a payback calculation, and the one most business cases rely on.

Where this argument is weakest

Three places, and the first would change the most if I am wrong.

Commissioned is not the same as wrong

Total Economic Impact studies publish their methodology, name their composite, state their discount rate and apply an explicit risk adjustment to benefits. That is more disclosure than almost any free industry survey offers.

The bias sits in which customers get interviewed and which studies reach publication, not in the arithmetic. A reader who discounts them entirely throws away the most detailed cost models available in public.

One trial is not a verdict on a function

METR is 16 developers, one tooling generation, and open-source repositories with unusually high review standards. METR says so itself, and warns against reading the result as evidence that AI fails developers generally.

Agent capability also moved between early 2025 and now. A 2026 rerun could land on the other side of zero, and I would not be surprised if it did.

The third weakness is mine. I could not find a primary source for the circulating by-function table, and that is not the same as proving none exists. A paywalled benchmark could carry those figures. What I can say is narrower: no public page I opened contains them with a methodology attached, and the pages repeating them disagree about what they measure.

How to build a payback number your CFO can check

The useful move is not to find a better benchmark. It is to build a small number you own, which beats a large number you borrowed.

Start with a unit of work that already carries a price in your business. Cost per support contact. Cost per invoice processed. Fully loaded cost per qualified meeting booked. If a function has no priced unit of work, it has no payback number either, and that is worth discovering before procurement rather than after.

What a vendor payback model counts, and what your model has to add
Cost lineIn the commissioned modelWhat to add before you sign
Licence and platform feesCounted, usually accuratelyModel the usage tier you land in at month 12, not month 1.
Implementation and integrationCounted as professional servicesInternal engineering time on data access, the barrier 41% of Bain respondents named first.
Change management and trainingSometimes counted, often thinHours per affected employee, times the number who actually change how they work.
Review and reworkRarely countedTime spent checking agent output, which is where the METR slowdown came from.
Incident and error handlingNot countedA budgeted allowance for the failures that reach a customer.
The counterfactualNever countedWhat the same budget and the same six months would return elsewhere.

The last row is the one finance teams ask about and the one no vendor study can answer, because a comparison against the next best use of the money is not a product benefit.

Record the baseline before the pilot starts, with a date on it. A baseline reconstructed afterwards gets reconstructed to fit, and the METR result shows how confidently people misremember their own throughput. The same discipline applies to the contract, where the clauses a CFO should require before signing an AI deal decide whether a disappointing month is renegotiable or simply invoiced.

Then price the thing properly. Agent pricing is usually consumption-shaped, so a model built on a flat licence assumption will be wrong by month nine. The conversation-by-conversation arithmetic behind Agentforce pricing shows how fast a per-unit rate compounds at real volumes. It is the small version of the question the hyperscalers face in the payback math behind 2026 AI capital spending, and the small version is the one you can answer.

Frequently asked questions

What is the average payback period for AI agents?

No trustworthy average exists. Six published Forrester Total Economic Impact studies report payback in under six months, but each is commissioned by the vendor and models a composite organisation built from four to six customer interviews. The one study using a 420-person survey reports 15 months. Independent buyer surveys do not publish a payback average at all, and only 13% of decision-makers report a positive EBITDA impact.

Which business function gets the fastest ROI from AI agents?

The published evidence cannot answer this. Studies covering customer service, contact centre, engineering, finance, IT operations and marketing all report payback in the same under-six-months band, so the data has no power to rank functions. Circulating tables that rank sales fastest and finance slowest do not trace to a primary source. Judge by whether the function has a priced unit of work instead.

Are Forrester Total Economic Impact studies reliable?

They are transparent and sponsored at the same time. Each publishes its methodology, composite organisation, discount rate and risk adjustments, which is more disclosure than most free surveys offer. The bias is in selection: the interviewed customers are supplied by the vendor, and weak results are unlikely to reach publication. Treat them as detailed cost models rather than as measurements of a market.

Why do AI agent ROI figures vary so much?

Because the inputs are chosen rather than observed. Reported ROI in the seven studies ranges from 120% to 449%, driven by the size of the modelled composite, which benefits are counted and how aggressively they are risk-adjusted. Scope matters most: single-function point solutions report the highest returns, while the one platform-wide study reports the lowest at 120%.

How do I calculate payback for an AI agent project?

Pick a unit of work that already has a price, such as cost per support contact or per invoice processed. Record the baseline before deployment, with a date. Add licence fees at the usage tier you expect by month 12, internal integration engineering, training hours, review and rework time, and an allowance for incidents. Then compare the result against the next best use of the same money.

Where to start this week

Take the agent business case currently sitting in your pipeline and do two things to it.

First, find every payback figure in it and write the primary source next to each one, with the sample size and who paid for the study. Anything you cannot trace to a document you opened yourself comes out of the deck. On current evidence that removes most of the numbers, which is the point.

Second, write down the unit of work the agent is supposed to make cheaper, and today's price for one of them. If nobody in the room knows that price, you do not have a business case yet. You have a procurement request with a chart on it.

Related analysis

For the wider version of this question, see where measurable AI return has actually shown up, and the six failure modes that stop agent pilots reaching production.

References

  1. Forrester, Three Questions That Will Define AI In 2026, 26 January 2026. Survey of over 1,400 global AI decision-makers. Used for the EBITDA, P&L and payback horizon figures.
  2. Bain & Company, Your AI Budget Is Growing. Your Returns Aren't., 1 June 2026. Automation and AI Pathfinder Survey 2026, 951 companies. Used for cost savings, autonomy and data barriers.
  3. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. Used for the 19% slowdown and the expectation gap.
  4. Forrester Consulting, The Total Economic Impact of Microsoft's Agentic AI Solutions, January 2026. Used for the 120% ROI and 15-month payback.
  5. Forrester Consulting, The Total Economic Impact of Salesforce Agentforce, November 2025. Used for the customer service row.
  6. Forrester Consulting, The Total Economic Impact of Google Cloud Customer Engagement Suite, May 2025. Used for the contact centre row.
  7. GitLab, Study finds organizations can achieve 400% ROI with GitLab Duo Agent Platform, 16 July 2026. Used for the engineering row.
  8. Vendor summaries of commissioned Forrester studies for LogicMonitor Edwin AI, FloQast and Insider One, 2026. Used for the IT operations, finance and marketing rows.

Weakest thing about this source base: every payback figure in the table comes from research a vendor paid for, and three rows were read from vendor summaries rather than the full study documents. The two independent sources here publish no payback-by-function data at all.

SK
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading