From Sanskriti Khandelwal | Product & Market Analysis

How Vendors Inflate AI ROI Case Studies, and the 7 Checks That Catch It

On this page

Most AI case studies are inflated without containing a single false number. The technique is selection: which baseline, which cohort, which window, and which costs stay off the page. A randomised trial run by METR shows how much room that leaves. Sixteen experienced developers were measured 19% slower when allowed to use AI tools, and the same developers reported afterwards that AI had made them 20% faster.

Key takeaways

  • The figures in vendor case studies are usually accurate. The sentence built around them is where the inflation lives. METR measured 16 developers across 246 real tasks at 19% slower, against a self-reported 20% speedup.
  • The most quoted AI productivity statistic describes one task. GitHub's 55.8% speed gain came from 95 developers writing an HTTP server in JavaScript, with a confidence interval running from 21% to 89%.
  • Self-reported return is a satisfaction survey with a currency sign in front of it. The widely cited $3.7 back per $1 spent comes from an IDC InfoBrief sponsored by Microsoft, drawn from over 4,000 business leaders describing their own results.
  • A case study is a snapshot, and the sequel is where the truth sits. Klarna announced work equivalent to 700 agents in February 2024, then began rehiring humans in May 2025 because quality had fallen.
19%slower with AI tools, measured, while the same developers reported a 20% speedup. Source: METR, July 2025.
55.8%faster on one JavaScript task, n=95. The figure most often quoted as general AI productivity. Source: GitHub, 2022.
$3.7returned per $1 of generative AI spend, self-reported by 4,000+ leaders. Source: Microsoft and IDC, November 2024.

What inflation in an AI case study actually is

An inflated AI ROI case study rarely carries a fabricated figure. It selects instead: a baseline captured at the worst possible moment, a cohort that already succeeded, a window short enough to annualise, and a cost line that omits vendor engineers and internal rework. Seven checks, run in order, catch most of it inside an hour.

This matters more in AI than in software generally, for one structural reason. Nobody has agreed what the unit of output is. A support ticket, a resolved conversation and a merged pull request are all countable, and none of them is the thing your business sells.

The number is usually real, which is what makes it work

Outright fabrication is rare in enterprise marketing, and the reason is exposure rather than virtue. The Federal Trade Commission opened five enforcement actions in September 2024 under the banner Operation AI Comply, including one against DoNotPay, which marketed itself as "the world's first robot lawyer" without substantiation. Then chair Lina Khan put the standard plainly: there is no AI exemption from the laws on the books.

The endorsement guides at 16 CFR Part 255 push the same way. If a testimonial shows results the advertiser cannot show are typical, the advertisement has to state what people can generally expect. So assume the number happened, then ask what conditions produced it and whether any of them will hold at your company.

The seven checks, in the order that saves the most time

Run them in sequence. The first two eliminate most weak case studies before you have spent twenty minutes, which is the point of ordering them this way.

Case study forensics: what to ask, and what a weak answer sounds like
CheckWhat you ask forA weak answer
1. BaselineThe pre-deployment number, when it was captured, and by whom"Compared to their previous manual process"
2. DenominatorThe full population, including the cases the system refusedA percentage with no total attached
3. EscalationsWhere deflected work went, and what it cost thereDeflection rate quoted as resolution rate
4. SupportVendor engineering hours during the pilot, and after it"Our team worked closely with them"
5. SourceWhether it was measured in a system or reported in a survey"Customers tell us they see..."
6. SubjectWhether the customer is one company or a modelled composite"A composite organisation based on interviews"
7. WindowLength of the measured period before any annualisationAn annual saving from a one-month trial

Checks 1 and 2: the baseline, and the denominator

A baseline is a number recorded before the purchase decision. Almost everything presented as a baseline was reconstructed afterwards by the team that had already chosen to buy. That is not fraud, it is memory, and memory favours the version that justifies the decision.

Who chose the baseline, and when

The clearest published example is also the most cited statistic in the category. In 2022, GitHub ran a controlled experiment with 95 professional developers, split randomly, on a single task: write an HTTP server in JavaScript. The Copilot group finished in 1 hour 11 minutes against 2 hours 41 minutes, a 55.8% gain, significant at P=.0017.

The study is honest and says what it did. The inflation happens downstream, when a figure from one greenfield task with a known solution is quoted as general productivity. GitHub also reported a confidence interval from 21% to 89%, which almost nobody repeats.

Now compare the design. METR's July 2025 trial gave 16 developers 246 real issues in repositories they had worked in for years, averaging over 1 million lines of code. They were measured 19% slower. They had forecast a 24% speedup, and afterwards they still believed they had gained 20%.

My position is that these two studies do not contradict each other, and treating them as rival headlines is the error. One measured a clean task, the other measured messy work in familiar code. Your engineers live in the second condition, so weight it accordingly, and note the same tension in the three scoreboards used to judge coding tools.

Three answers to the same question, from the same 16 developers METR randomised trial, 246 real tasks in repositories they already knew. Bars are percent change in speed. no change Predicted before +24% faster Reported after +20% faster Measured 19% slower Source: METR, 10 July 2025.
The two blue bars are what a case study collects. The red bar is what a stopwatch collects. Both came from the same people in the same week.

A denominator you can actually see

A percentage without a denominator is a decoration. "Resolves 70% of tickets" tells you nothing until you know whether the 70% is of all inbound contacts, of contacts the system agreed to accept, or of one narrow category.

Good disclosure exists, which is why the standard is fair to demand. Intercom's Fin publishes benchmarks drawn from more than 110 million conversations across 12,000 customers, defines a resolution as the customer confirming the issue is solved or not following up, and states that the headline figures are averages of the top 10 performers in each industry. That last sentence is the one most vendors omit. Its pricing consequences are worked through in the breakdown of resolution-based pricing.

Ask for the denominator in writing. A vendor who can produce it has an instrumented product. A vendor who cannot has a dashboard.

Check 3: count the work that moved rather than disappeared

Deflection and resolution are different events, and the gap between them is where support automation savings go. Deflected work reappears as repeat contacts, as escalations into a smaller and more expensive senior queue, and sometimes as churn. None of those lines sit in the case study, because none of them belong to the team that ran the pilot.

Klarna published both halves, 15 months apart

Klarna's February 2024 announcement is the most reproduced AI support case study in existence, and the precision of it deserves credit. In its first month the assistant handled 2.3 million conversations, two-thirds of chats, described as the equivalent work of 700 full-time agents. Resolution time fell from 11 minutes to under 2, and repeat inquiries dropped 25%.

Read the profit line carefully. The release says the assistant was "estimated to drive" a $40 million profit improvement in 2024. That is a projection about a first month, and the trade press reproduced it as an achieved saving.

In May 2025, chief executive Sebastian Siemiatkowski told Bloomberg that Klarna was hiring human agents again, because cost had been too dominant a factor and "what you end up having is lower quality". Nothing in the 2024 figures was false. They simply did not measure the thing that later forced the reversal, and I would treat any support case study without a 12-month follow-up as a first-month result. The later failure modes are catalogued in the piece on how agent pilots actually fail.

The case study and its sequel Klarna AI customer service assistant, as described by the company itself Feb 2024 2.3M conversations two-thirds of chats "equivalent of 700 agents" 2024 estimate $40M profit improvement stated as an estimate, repeated as a result May 2025 rehiring human agents CEO: "what you end up having is lower quality" Sources: Klarna press release, 27 February 2024; Siemiatkowski to Bloomberg, 8 May 2025.
Notice what is missing between the second and third markers. Nobody published a measurement of quality during the 15 months in which quality was the thing that changed.

Check 4: the vendor labour that never reaches the invoice

Every impressive AI pilot has a staffing footnote. Someone from the vendor mapped the data, wrote the prompts, tuned the retrieval and fixed the integration that broke in week three. That work is real, it is expensive, and it is usually free during an evaluation.

The category now has a job title. TechCrunch reported in July 2026 that forward-deployed engineering has become the industry's dominant hiring obsession, citing a Christian and Timbers study of more than 250 executives in which the share of companies planning to hire such engineers rose from between 5% and 10% at the start of 2026 to 70% by the second quarter. Read that as a cost disclosure, because that is what it is.

Ask what month four looks like

The most useful question in an AI reference call is not about results. It is: how many hours a week did the vendor's staff spend on this in month one, and how many in month four?

A steep decline means the product carried the load. A flat line means you are buying a consulting engagement with a licence attached, which can be a fine purchase as long as you price it as one. The distance between a demonstration and a system that survives without hand-holding is the subject of the production readiness tests worth running before you sign.

Checks 5 and 6: measured, reported, and modelled

Three different objects arrive in the same font. A measured figure comes out of a system, a reported figure comes out of a person, and a modelled figure comes out of a spreadsheet built from several people. All three can be legitimate, and they carry different weight.

Self-reported return is a satisfaction survey with a currency sign

The most repeated ROI figure in enterprise AI is $3.7 returned for every $1 spent, rising to $10.3 for leading adopters. It comes from an IDC InfoBrief sponsored by Microsoft, published November 2024, drawn from more than 4,000 business leaders and AI decision-makers.

Nothing about that is improper, and IDC discloses the sponsorship. It is a survey of what executives believe about their own returns. METR is the reason the distinction is not pedantic: belief and measurement, collected from the same people about the same work, pointed in opposite directions by 39 percentage points.

The measured picture looks different. Research summarised in the analysis of where generative AI return has actually appeared found most enterprise pilots producing no measurable profit and loss impact, which is hard to reconcile with a broad 3.7 times return.

A composite organisation is not a customer

Vendor-commissioned economic impact studies are the most rigorous marketing artefact in B2B software, and they are still marketing artefacts. Forrester's Total Economic Impact method interviews real customers, builds a single "composite organisation" from them, then models costs and benefits over three years with a risk adjustment applied. Forrester states that all such studies are marked as commissioned and cannot appear in its syndicated research.

The disclosure is the good part. The trap is reading a composite as a customer. A composite organisation is an average with a headcount, and averages do not have staff turnover or a data migration that slipped two quarters.

Forrester's own analysts are sceptical of the wider genre. Sam Higgins wrote in August 2025 that big technology economic impact claims are rarely revisited, with no formal tracking of whether the promised benefits materialise. He was writing about national economic impact claims rather than about TEI, and the transferable point is the one worth keeping: nobody goes back and checks.

What each kind of evidence actually establishes Ranked by how much of the result survives a change of context. Bar length is that judgement, not a measurement. Randomised trial, real tasks Causal, in the tested setting Instrumented product data Real, if the denominator is shown Named single-customer study One organisation, one context Modelled composite study An average with a headcount Self-reported survey What buyers believe Most headline AI ROI numbers in circulation sit in the bottom two rows.
This ranking is illustrative rather than measured. The ordering is a judgement about transferability, and the bar lengths carry no units.

Check 7: the window, and what annualising it assumes

Short windows flatter AI deployments, and the reason is mechanical. Easy cases arrive first because teams route them there deliberately, so the early sample is not representative of steady state. Annualising a first month multiplies that selection effect by twelve, and novelty effects push the same way.

The honest counter is that some benefits take longer than a quarter to appear, so a short window can understate as well. That is exactly why the window belongs in the case study rather than in your assumptions, and why I would rather see a modest 12-month number than a spectacular 30-day one.

What each measurement window can hide
WindowWhat it flattersWhat it cannot yet show
30 daysNovelty usage, easiest cases, vendor staffing at its peakRepeat contacts, escalation cost, quality drift
One quarterAdoption curve during active vendor supportRenewal pricing changes, model or packaging revisions
12 monthsLittle. This is the first window that behaves like realitySecond-year price increases and switching cost

The renewal is the real test, and most case studies were never designed to survive it. Prices move, usage allowances tighten, and the champion who ran the pilot has often changed jobs. That repricing is set out in the piece on what happens to grandfathered pricing at renewal.

Where this argument is weakest

A forensics checklist can turn into a licence to dismiss any evidence you did not generate yourself. That is a worse error than believing a case study, so here is where the argument gives way.

Buyers inflate at least as much as vendors do

The METR result cuts both ways. The developers overstating their gains were not selling anything. They reported honestly on their own experience and were wrong by 39 percentage points, which suggests the problem is human measurement rather than vendor ethics. Before interrogating a vendor's baseline, check whether you recorded your own, a habit whose absence shows up in the analysis of deployments that returned less than they cost.

There is also a scarcity argument in the vendor's favour. Almost nobody has a clean 12-month controlled measurement, because these products change faster than a measurement period. Demanding one can mean rejecting every option and keeping a status quo that was never measured either. Buyers who prefer verified outcomes to vendor narrative should read what actually replaces the B2B case study before discarding the format.

Frequently asked questions

How do you verify a vendor's AI ROI case study?

Ask for four things in writing: the baseline figure with the date it was captured, the denominator behind every percentage, the vendor engineering hours spent during and after the pilot, and the length of the measured window before any annualisation. A vendor with an instrumented product can supply all four within a week. One who answers with a multiplier or a testimonial has not measured what the case study claims.

Why do AI case studies show better results than internal pilots?

Three reasons, and only one of them is marketing. Published deployments are selected from successes, so failures never reach the page. They usually ran with vendor engineers embedded, which you may not get. And they were measured in a short early window when the easiest cases were routed to the system first. The same deployment measured at 12 months routinely looks weaker than it did at 30 days.

What is a good baseline for measuring AI productivity?

A number recorded before anyone decided to buy, taken from a system rather than from memory, covering a full business cycle. Cycle time, cost per resolved case, rework rate and escalation rate all work. Any baseline reconstructed after the purchase decision is biased toward justifying it. If you have no pre-purchase figure, record one now and delay the deployment by a month rather than measuring against a guess.

Are vendor-commissioned ROI studies trustworthy?

They are useful and they are not independent. Forrester's Total Economic Impact studies interview real customers, model a composite organisation over three years, apply a risk adjustment, and are labelled as commissioned. Read the composite as a model rather than as a customer. Treat the benefit categories as a checklist of what to measure yourself, and treat the totals as the vendor's best case with a methodology attached.

What questions should you ask a vendor about a case study?

Ask who chose the baseline and when. Ask what the percentage is a percentage of. Ask where deflected work went and what it cost there. Ask how many vendor hours went in during month one and during month four. Ask whether the figure was measured in a system or reported in a survey. Then ask to speak to a customer at 18 months rather than at launch.

Is it illegal for a vendor to publish an inflated AI case study?

Unsubstantiated performance claims can be deceptive under Section 5 of the FTC Act, and the same standard applies to business buyers as to consumers. The FTC brought five actions in September 2024 under Operation AI Comply, including one against a firm marketing a "robot lawyer" it could not substantiate. Most enterprise case studies sit well short of that line, because selection and framing are not the same thing as a false statement.

Where to start this week

Take the last AI case study a vendor sent you and run checks 1 and 2 on it. Put the two questions in an email: what was the baseline, captured when, and what is the denominator behind the headline percentage. The reply time alone is informative.

Then do the harder half. Pick one AI tool already in your stack and find whether a baseline exists anywhere. If it does not, record one this week for whatever the tool was bought to change. You cannot audit a vendor's arithmetic credibly while carrying none of your own.

Related on buying discipline

If you are assembling a shortlist rather than checking one claim, the companion piece on what to demand during AI agent procurement covers the contract side of the same problem.

References

  1. METR, Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity, 10 July 2025. The 19% slowdown, the 24% forecast and the 20% post-task estimate.
  2. GitHub Blog, Research: quantifying GitHub Copilot's impact on developer productivity and happiness, 7 September 2022. The 55.8% figure, the task design and the confidence interval.
  3. Klarna, AI assistant handles two-thirds of customer service chats in its first month, 27 February 2024, and Entrepreneur, 9 May 2025, on Siemiatkowski's comments to Bloomberg about the reversal.
  4. Microsoft, IDC's 2024 AI opportunity study, 12 November 2024. The $3.7 and $10.3 self-reported returns, and the sample of 4,000+ leaders.
  5. Federal Trade Commission, FTC announces crackdown on deceptive AI claims and schemes, 25 September 2024. Operation AI Comply and the Khan quotation.
  6. Forrester, Total Economic Impact methodology, and Sam Higgins, Beyond the hype: why big tech economic impact studies fall short, 15 August 2025.
  7. Fin, public benchmarks, refreshed 20 May 2026. The resolution definition, the 110 million conversation base and the top-performers caveat.
  8. TechCrunch, Forward-deployed engineers are the AI industry's latest talent obsession, 30 July 2026. The Christian and Timbers hiring figures.

The weakest part of this source base is that both of the strongest exhibits, the METR trial and the GitHub experiment, measure software engineering. Whether the same perception gap holds in support, sales or finance work has not been tested with comparable rigour, so these checks generalise further than the evidence does.

SK
Sanskriti Khandelwal
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading