From Shubhi K | Product & Market Analysis
Who Pays When the AI Agent Gets It Wrong: Caps, Credits and Real Remedies
On this page
When an AI agent makes a costly mistake, the buyer usually pays. Standard software contracts cap the vendor's total liability at the fees you paid, and exclude the loss categories an agent is most likely to cause. Two real cases and one new insurance market show what that gap is worth.
Key takeaways
- The customer carries the risk by default. Clifford Chance's February 2026 briefing states it directly: the business procuring the technology bears the risk of actions taken by AI agents, because suppliers disclaim accuracy and fitness for purpose.
- The two remedies that have actually paid out were not contractual. A tribunal ordered Air Canada to pay C$650.88 over a chatbot answer. Deloitte refunded A$97,000 on a A$440,000 report. Neither came from a liability clause.
- Performance guarantees are appearing, and they are narrow. Intercom publishes a $1 million refund guarantee on Fin and a $1 million resolution-rate commitment, both capped, both conditional, both rare in the market.
- Insurers priced the gap before contracts closed it. Chaucer and Armilla launched $25 million of dedicated AI limits in February 2026, while several US carriers filed to exclude generative AI from general liability cover.
What "AI agent liability" actually means in a contract
Liability is not one clause. It is four separate mechanisms that most buyers treat as a single line item, and they fail in different ways.
The first is the cap, which sets the maximum the vendor can ever owe you. The second is the exclusions list, which removes whole categories of loss from recovery before the cap is even reached. The third is the indemnity, where the vendor agrees to defend you against a specific type of third-party claim. The fourth is the service level, which pays out in credits when availability slips.
Agent failures do not fit any of them cleanly. An agent that misprices an order, approves a payment or emails a customer the wrong commitment does not cause downtime. It causes a business decision that turns out to be wrong, and the loss lands in the categories the exclusions list already removed.
That mismatch is the whole subject. It is not a drafting oversight. It is the deliberate shape of software contracts, applied to a product that no longer behaves like software.
The three-link chain, and where it breaks
Almost no agent is a single product. A typical deployment has a model provider at the bottom, an application vendor in the middle, and your business on top, with your customer on the receiving end.
Each link has its own contract, its own cap and its own exclusions. Those documents were negotiated separately and none of them was drafted to be read alongside the others. The result is a chain where liability stops at each joint rather than passing through.
The model provider caps at what the vendor pays it, which is often a fraction of what you pay the vendor. The vendor caps at what you pay. You face your customer with no cap at all, because consumer and commercial law does not offer you one. The exposure is largest at exactly the point where the contractual protection is smallest.
What AI vendor contracts promise in 2026
Clifford Chance published a briefing in February 2026 on exactly this question, and its finding is blunt. The customer bears the risk of actions taken by AI agents, because suppliers provide the service on an as-is basis and disclaim accuracy, reliability and fitness for purpose.
The cap is a ceiling, not a promise
A cap does not create a right to be paid. It sets the maximum you could recover if you first establish that the vendor breached something. Most buyers read the number and stop there.
The common shape ties the cap to fees paid over the trailing period, often twelve months. On a $100,000 contract that is a $100,000 ceiling, regardless of what the agent actually cost you. Treat that number as the vendor's estimate of its own worst case, not yours.
The exclusions remove the losses agents cause
Technology contracts routinely exclude loss of profits, loss of data, and consequential, indirect and special damages. Clifford Chance's analysis of agentic deployments notes the mismatch: agent failures produce regulatory exposure, revenue disruption and reputational damage, which is the excluded list almost word for word.
This is the part I would fight before the cap. A generous cap over a wide exclusions list pays nothing. A modest cap with a carve-out for agent-caused loss pays something.
Indemnities cover copyright, not wrong decisions
The indemnities that vendors advertise are intellectual property indemnities. Microsoft's Copilot Copyright Commitment, OpenAI's Copyright Shield and Google's generative AI indemnity all defend you against third-party claims that an output infringed someone's copyright.
None of them covers an agent that authorises the wrong payment or gives a customer the wrong answer. Those are the failures your business will actually meet, and Clifford Chance notes that standard indemnities do not typically extend to an AI agent's acts or omissions.
| Failure the agent causes | Covered by a standard contract? | Why |
|---|---|---|
| Platform is down and the agent cannot respond. | Yes, partly. | Availability service levels pay credits or allow termination. |
| Output infringes a third party's copyright. | Yes, if you use the safety controls. | Vendor IP indemnities were written for this case. |
| Agent invents a policy and a customer relies on it. | No. | Accuracy is disclaimed, and the loss is usually consequential. |
| Agent approves an incorrect payment or refund. | No. | Direct financial loss, but capped at fees and often excluded. |
| Agent decision triggers a regulatory penalty. | No. | Fines sit in the excluded categories and compliance sits with you. |
Rows describe the market default set out in the Clifford Chance briefing, not any single named agreement. Negotiated enterprise contracts differ, and the differences are the point of this post.
The remedies buyers are actually getting
The picture is not uniformly bad. A small number of vendors now put money behind performance, and the mechanisms they use are worth copying into your own negotiations.
Performance guarantees exist, and they are conditional
Intercom is the clearest public example. Its published guarantee refunds new customers their full Fin spend up to $1 million if they are not satisfied within 90 days, and commits $1 million against a 65% resolution rate for enterprises above 250,000 monthly conversations.
Read the conditions, because they carry the weight. Both commitments are capped at $1 million. One expires after 90 days. The other applies only at a volume most buyers never reach. This is a real commercial commitment and it is not a liability regime.
The service level sits separately and is narrower still. Intercom's service level agreement targets 99.8% monthly uptime for the AI agent, and the stated remedy after two consecutive failed months is termination with a refund of prepaid fees. There are no service credits at all, and uptime measures whether Fin can generate a response, not whether the response was right.
Outcome pricing is a partial remedy
Charging per resolved conversation rather than per seat does transfer some risk. If the agent fails to resolve, you do not pay for that attempt. Zendesk and HubSpot both price this way, and we covered the arithmetic of the same model in the breakdown of Agentforce at $2 a conversation.
The limit is easy to miss. Outcome pricing refunds the price of the failure, not the cost of it. A $2 resolution that gives a customer a commitment you then have to honour has cost you far more than $2, and the pricing model returns the $2. It is a discount mechanism wearing the clothes of a guarantee.
That distinction matters more as pricing shifts away from seats, a move traced in the autopsy of per-seat pricing. Buyers are being offered outcome alignment as a substitute for accountability. They are not the same purchase.
| Mechanism | What it pays | Where it fails |
|---|---|---|
| Availability credits. | A share of monthly fees. | Measures whether the agent responded, not whether it was correct. |
| Prepaid fee refund on termination. | Unused fees only. | Requires sustained failure, and returns nothing for a single costly error. |
| Accuracy or resolution guarantee. | A fixed capped sum. | Conditional on volume, tenure or a measurement the vendor controls. |
| Negotiated liability carve-out. | Above-cap recovery for named agent failures. | Rare, and it costs commercial ground elsewhere in the deal. |
| Third-party AI liability insurance. | Limits independent of the vendor contract. | You pay the premium, and cover depends on governance evidence. |
Two cases that show who actually pays
Almost every discussion of this topic runs on hypotheticals. There are two documented outcomes worth more than all of them, and both point the same way.
In February 2024 the British Columbia Civil Resolution Tribunal ordered Air Canada to pay C$650.88 to a passenger whose chatbot answer described a bereavement refund policy that did not apply. The airline argued the chatbot was a separate entity responsible for its own answers. The tribunal rejected that outright and held the airline responsible for everything on its own website.
The sum is trivial. The allocation is not. The operator paid, the operator's technology supplier paid nothing, and the reasoning generalises to any agent your customers can reach.
The second case runs the other way, and it is the more useful one. Deloitte delivered a compliance review to Australia's Department of Employment and Workplace Relations for close to A$440,000. After a University of Sydney academic identified citations to research that did not exist, Deloitte repaid over A$97,000, the final instalment.
That refund is roughly 22% of the contract. It was not compelled by a liability clause, a service level or a court. It was a commercial decision made by a supplier that wanted to keep a client, which is the most common remedy in this category and the least reliable one.
Insurance repriced the gap before the contracts did
The fastest signal on this question is not coming from legal departments. It is coming from underwriters, who have to put a number on the risk whether or not the contracts are ready.
In November 2025 the Financial Times reported that several large US carriers, including Great American, Chubb and W.R. Berkley, had asked regulators for permission to exclude generative AI liabilities from general liability policies. As TechCrunch summarised the reporting, the concern was not a single large claim but correlation: one agent fault repeating across thousands of customers at once. AIG later said it had no plans to implement the exclusions it was reported to have sought.
Three months later the opposite trade appeared. Chaucer and Armilla AI launched Vanguard AI on 10 February 2026, pairing $10 million of cyber limits with $25 million or more of dedicated AI liability limits per organisation, underwritten at Lloyd's. The covered exposures are named explicitly: erroneous outputs, model underperformance and AI agent actions, including cases where no cyber event has occurred.
Read those two developments together. General liability is being narrowed and specialist AI cover is being sold at a premium into the space it vacates. That is a market saying the risk is real, priceable and currently sitting in the wrong place.
The law is moving in the same direction, slowly. The European Commission withdrew its proposed AI Liability Directive in February 2025, as Bird & Bird recorded at the time. What remains is the revised Product Liability Directive, which treats software and AI systems as products under strict liability and must be transposed by member states in December 2026. Strict liability removes the need to prove negligence, which changes the calculus for anyone selling an agent into Europe.
Where this argument is weakest
The buyer-side case above is the one I hold. It has three genuine soft spots, and a post that hid them would fail its own standard.
The case for the vendor's cap
A vendor charging $60,000 a year cannot underwrite an unbounded loss on a customer's revenue. If it accepted that exposure across a thousand customers, it would either raise prices to cover the tail or fail on the first claim.
Caps are also what make software cheap. The alternative is a professional services model with insurance loaded into the price, and most buyers would decline that trade if it were priced honestly in front of them.
What two cases cannot tell you
Air Canada and Deloitte are two data points, one of them a small claims matter and neither of them a contract dispute between an enterprise and an AI vendor. They show where losses landed. They do not establish a norm, and treating them as one would be exactly the error this publication warns about elsewhere.
The honest position is that the contractual norms are still forming. Negotiated enterprise agreements are private, so nobody outside the parties can measure how often carve-outs are actually granted. Anyone quoting a percentage for that is guessing.
The clauses worth your negotiating capital
You will not win every point, so spend the capital where the payout is largest. My ranking below puts evidence rights above money, which is the opposite of most buyer checklists.
| Ask for this | Why it ranks here | Likely vendor response |
|---|---|---|
| Decision logs and an explainability right. | Without evidence you cannot prove causation, so every other clause is unusable. | Often granted, because it costs the vendor little. |
| A real-time suspension right. | Caps the size of a repeating fault rather than arguing about it later. | Negotiable, usually with notice conditions. |
| Carve-out from the exclusions for agent-caused direct loss. | Restores the losses the standard list removes. | Resisted hardest, and the truest test of vendor confidence. |
| A supercap for agent failures, above the general cap. | Money, but only after the two rows above exist. | Sometimes granted at a multiple of the general cap. |
| Evidence of the vendor's own AI liability cover. | Tells you whether anyone has underwritten this product. | Increasingly available, and revealing when it is not. |
One test cuts through the whole negotiation. Ask the vendor what it has paid out under these terms in the last twelve months, with a count and a total. A vendor that has paid something has a working process. A vendor that answers with reassurance has an untested clause, which is the same thing as no clause.
Where you are choosing between building and buying, this changes the sum. Building keeps every failure on your own balance sheet, but so does buying, and we set out the rest of that comparison in the build versus buy analysis for coding agents. The liability difference between the two options is smaller than most buyers assume.
The chain problem gets worse as agents call other agents through shared protocols, a direction covered in the piece on MCP as an interoperability standard and in the analysis of agent marketplaces. Every new link adds a cap and an exclusions list you did not negotiate.
Frequently asked questions
Who is liable when an AI agent makes a mistake?
In practice the business deploying the agent is liable to the affected party. Clifford Chance's February 2026 briefing states that the customer procuring the technology bears the risk of actions taken by AI agents. The British Columbia Civil Resolution Tribunal reached the same result in Moffatt v. Air Canada, holding the airline responsible for its chatbot's answer and rejecting the argument that the chatbot was a separate entity.
Do AI vendors cap their liability for agent errors?
Almost always. The standard structure caps total liability at the fees the customer has paid, commonly measured over the trailing twelve months, and then excludes loss of profits, loss of data and consequential damages before the cap applies. Because agent failures usually produce exactly those excluded categories, the practical recovery is often far below the headline cap number written into the contract.
What is the difference between a service credit and a refund in an AI contract?
A service credit is a discount against future fees when availability falls below target. A refund returns money already paid. Intercom's published service level agreement offers neither credits nor partial refunds for poor answers: it targets 99.8% monthly uptime and, after two consecutive failed months, allows termination with a refund of prepaid fees. Uptime measures whether the agent responded, not whether it was right.
Does AI vendor indemnification cover agent errors?
Usually not. The widely advertised commitments from Microsoft, OpenAI and Google are intellectual property indemnities covering third-party copyright claims about generated output, and they require you to keep the built-in safety controls switched on. They do not cover an agent that approves the wrong payment, misprices an order or gives a customer an answer your business then has to honour.
Can you insure against AI agent errors?
Yes, and the market grew quickly through 2026. Chaucer and Armilla launched Vanguard AI in February 2026 with $25 million or more of dedicated AI liability limits per organisation, covering erroneous outputs, model underperformance and AI agent actions. The same period saw several general liability carriers file to exclude generative AI risks, so specialist cover is filling a space that standard policies are leaving.
What should an AI agent SLA include?
Availability alone is not enough. Ask for decision logs and an explainability right so you can prove what happened, a suspension right so a repeating fault can be stopped in real time, and a defined accuracy or resolution measure with a stated method. Then negotiate money: a carve-out from the exclusions for agent-caused direct loss, and a supercap above the general liability cap.
Where to start this week
Pull the contract for the one agent that touches your customers or your money, and find three things in it. The cap, the exclusions list, and any clause giving you access to logs. If the third does not exist, that is the finding, and it is the cheapest thing to fix at renewal.
Then price the exposure honestly. Write down the largest single loss that agent could cause in one day, and compare it to the annual fee. If the first number is a multiple of the second, you are self-insuring a risk you never chose to underwrite, and that decision belongs in front of whoever signs the renewal.
Related on pricing and accountability
If your agent bills by outcome, the pricing question and the liability question are the same question. See the Agentforce conversation math and where measurable AI return has actually shown up.
References
- Clifford Chance, Agentic AI: the liability gap your contracts may not cover, February 2026. Used for the allocation of risk, exclusions and indemnity scope.
- McCarthy Tetrault, Moffatt v. Air Canada: a misrepresentation by an AI chatbot, 2024. Used for the C$650.88 award and the tribunal's reasoning.
- CFO Dive, Deloitte refunds over $60K for report with AI errors, October 2025. Used for the A$440,000 contract value and the A$97,000 refund.
- Intercom, Service Level Agreement, effective 11 July 2025. Used for the 99.8% uptime target and the termination remedy.
- Fin by Intercom, AI agent pricing comparison, 2026. Vendor source, used only for that vendor's own published guarantee terms.
- Chaucer Group, Chaucer and Armilla AI launch Vanguard AI, 10 February 2026. Used for the $25 million AI limits and the covered exposures.
- TechCrunch, AI is too risky to insure, say people whose job is insuring risk, 23 November 2025, reporting Financial Times findings. Used for the carrier exclusion filings.
- Bird & Bird, Proposed EU AI liability rules withdrawn, 2025. Used for the withdrawal of the AI Liability Directive and the position of the revised Product Liability Directive.
The weakest part of this source base is the absence of negotiated enterprise agreements, which are private. Every statement about market-standard terms rests on law firm description and published vendor documents, not on a sample of signed contracts. Insurance limits and vendor guarantee terms change without notice, and both were checked on 21 August 2026.
Related reading