From Aryan Vatsa | Product & Market Analysis
AI Provider Contracts Compared: The Four Terms That Decide Your Risk
On this page
Buyers compare model providers on price per million tokens. The four terms that decide your exposure are data retention, training rights, uptime remedies and IP indemnity, and they differ more than the prices do. Anthropic publishes a liability cap of 12 months of fees. Google trains on free-tier API content and does not train on paid. Neither lab backs its first-party API with a credit-bearing uptime promise in its published terms.
Key takeaways
- Retention defaults look identical and are not. Anthropic deletes API inputs and outputs within 30 days. OpenAI keeps abuse-monitoring logs up to 30 days. Google logs paid Gemini prompts for a limited period it does not put a number on.
- Zero data retention is a configuration project, not a checkbox. Anthropic excludes batch, code execution and stateful agents. OpenAI excludes the assistants, threads and conversations endpoints. Google requires caching off and an abuse-monitoring exception.
- The free tier is where training rights actually differ. Google's terms say unpaid Gemini API content improves its products, and that human reviewers may read it. Paid usage is excluded from both.
- A court order beat a retention policy in 2025. OpenAI was ordered to preserve output logs it would otherwise have deleted, and zero-retention API customers were the ones the order could not reach.
The four terms that decide your risk
Which provider is safer to sign with depends on four clauses, not on price. Retention sets how long your prompts exist. Training rights set whether they improve someone else's model. The uptime clause sets whether an outage costs the vendor anything. Indemnity sets who pays when an output draws a claim.
This is written for the founder or director who signs the agreement, because that person owns the residual risk when a clause turns out to be thin.
| Term | Anthropic | OpenAI | |
|---|---|---|---|
| Default retention on the API | Deleted within 30 days | Abuse logs up to 30 days | Logged for a limited period, no number given |
| Trains on your content | No, under commercial terms | No, since 1 March 2023, unless you opt in | No on paid. Yes on unpaid, with human review |
| Zero retention path | By request, per organisation, then per feature | By approval, per eligible endpoint | By configuration, plus an abuse-monitoring exception |
| Output IP indemnity | Yes, with six carve-outs | Yes, conditional on using safety features | Yes, split into training data and output |
Sources are each provider's own terms and documentation, listed at the end. The Google row describes the Gemini API terms and Google Cloud commitments, not the consumer Gemini app policy.
Data retention: three defaults that only look the same
Retention is the clause security review reads first, and the one vendors describe most loosely. All three land near 30 days. What sits inside those 30 days differs.
Anthropic deletes in 30 days, then the exceptions start
Anthropic's commercial policy is specific. It automatically deletes API inputs and outputs within 30 days of receipt or generation. Four exceptions apply.
Flagged content is kept far longer. Anthropic retains flagged inputs and outputs for up to 2 years, and trust and safety classification scores for up to 7 years. Its compliance activity feed and remote session transcripts are retained for 6 years by default.
That last number surprises people. A team can hold a 30-day belief in its data map while agent session transcripts sit under a six-year policy.
OpenAI keeps abuse logs, not training data
OpenAI's default is structurally similar. API inputs and outputs are retained up to 30 days for abuse monitoring, then removed unless the law requires otherwise. Its documentation is blunt about training: data sent to the API has not been used to train or improve models since 1 March 2023 unless you explicitly opt in.
The gap is not the 30 days. It is that abuse-monitoring retention is a service practice rather than a contractual deletion guarantee, a distinction a court tested in 2025.
Google's line runs between paid and unpaid
Google draws its line between paid and unpaid use of the same API, not between products. The Gemini API additional terms say that for paid services Google does not use your prompts or responses to improve its products, and logs them for a limited period solely to detect prohibited use.
Two specifics matter. Grounding with Google Search stores prompts for 30 days. And on Google Cloud, the data processing addendum commits to deleting customer data after termination within a maximum period of 180 days, after a recovery period of up to 30 days.
That 180-day bound covers all customer data on the platform, not just prompts. It is not comparable to a 30-day prompt deletion promise, and both sides will quote whichever number flatters them.
Zero retention is a configuration project, not a checkbox
Every provider sells zero data retention as the answer. Every provider then excludes a list of features from it, and the exclusions land on the features teams adopt second.
What zero retention does not cover
Anthropic's documentation is the most detailed of the three, which is to its credit. Zero retention is enabled per organisation and applies per feature. Its eligibility table marks batch processing ineligible with 29-day retention, code execution ineligible with container data held up to 30 days, and managed agents ineligible because sessions are stateful and persist until you delete them.
A model-level rule catches people too. Anthropic designates Claude Fable 5 and Claude Mythos 5 as Covered Models requiring 30-day retention, so zero retention is unavailable for them. A zero-retention organisation calling those models gets a 400 error until it enables retention for a specific workspace.
OpenAI runs the same idea at endpoint level. Chat completions, responses, embeddings, images, audio and moderations are eligible. Assistants, threads and conversations are not. Google's route differs again: disable data caching per project, then opt out of abuse-monitoring prompt logging, which for most customers means a form or invoiced billing.
| Provider | Granularity | Notable exclusions |
|---|---|---|
| Anthropic | Per organisation, then per feature | Batch, code execution, managed agents, Console and Playground, Covered Models |
| OpenAI | Per endpoint, pre-approved customers | Assistants, threads and conversations endpoints |
| Per project | Requires caching disabled and an abuse-monitoring exception. Some advanced features are excluded |
The Google row is the least precise of the three because Google's data governance pages render client side and could not be read directly for this piece. Read it as the shape of the requirement, not the full list.
One consequence is for engineers rather than lawyers. Under Anthropic zero retention, cross-origin browser calls are not supported, so requests must route through a backend proxy. A legal setting arrives as a build decision, the same pattern described in the piece on how the agent interoperability standard changes integration work.
Training rights: the free tier is the whole difference
On paid commercial terms, all three say the same thing. Anthropic's terms state that Anthropic may not train models on customer content from the Services, and that the customer retains rights to inputs and owns outputs. OpenAI stopped training on API data in March 2023. Google commits to no training on paid prompts and responses.
The difference appears the moment somebody uses an unpaid key. Google's terms say that for unpaid services it uses submitted content and generated responses to improve its products and machine learning technologies. Human reviewers may read, annotate and process API input and output, disconnected from the account first.
Read that twice if your team prototypes on a free key. It is published, not hidden, and it is the most common way commercial data ends up outside the commercial agreement. I would not allow a free-tier key anywhere near customer data, and that is a five-minute policy to enforce at the billing account level.
Which account signs up matters too. Anthropic's zero-retention arrangement attaches to one organisation and does not extend to others under the same account.
What a court did to one provider's retention policy
Retention policies describe intent. Litigation describes what happens. In 2025 the difference was demonstrated at scale.
On 13 May 2025, in the New York Times copyright litigation, a magistrate judge ordered OpenAI to preserve and segregate all output log data that would otherwise have been deleted. That included data users had asked to delete. OpenAI appealed, arguing the order conflicted with its own privacy commitments.
What was carved out, and why it mattered
The carve-outs are the useful part for a buyer. Zero data retention API traffic was unaffected, because nothing was stored to preserve. ChatGPT Enterprise sat outside the requirement. Consumer accounts carried the exposure.
The order was later narrowed. A further order freed OpenAI from preserving new logs as of 26 September 2025, while data already preserved stayed available and flagged accounts remained on hold. On 7 November 2025 the court ordered production of 20 million de-identified consumer conversations, and denied a stay on 13 November 2025.
The lesson generalises, and no provider can contract around it. A retention promise is only as strong as the tier you bought and the litigation your vendor is in. The strongest position is not a well-worded clause. It is data that was never stored.
Uptime: the labs publish dashboards, the clouds publish remedies
Availability is where the direct relationship is weakest, and where the comparison favours buying the same model somewhere else.
Anthropic's published commercial terms contain no uptime percentage, no service credit and no availability remedy. I read the document looking for one. Any commitment from the labs is negotiated into an enterprise order form, which makes it invisible to you until you are already in a sales process.
Contrast a cloud reselling the identical models. Amazon's Bedrock service level agreement, effective 4 October 2023, publishes a credit table that applies to every customer by default. Same Claude model, remedy attached, nothing to negotiate.
What a credit table actually pays
Do not overrate the remedy. Bedrock pays a 10% credit for monthly uptime between 99.0% and 99.9%, 25% between 95.0% and 99.0%, and 100% below 95.0%. The claim is on you: filed within two billing cycles, with per-interval evidence, and credits under one dollar are not issued.
A 10% credit will not cover what an outage costs your business. The value is not the money. A uniform published remedy is a signal about operational maturity, and it gives procurement something concrete to hold. Who absorbs that cost is the subject of the piece on what inference actually costs the people serving it.
Indemnity: who defends you when an output is challenged
All three now indemnify business customers for IP claims arising from outputs. The differences sit in the carve-outs, and the carve-outs are where a real claim would land.
The three indemnities, read closely
Anthropic's Section K commits it to defend a customer against third-party claims that paid, authorised use of the Services or Outputs violates IP rights. Six carve-outs follow. The first three are modifications you made, combination with technology not provided by Anthropic, and your own inputs and data. The rest are use you knew was infringing, the practice of a patented invention contained in an output, and trademark claims from using an output in trade.
OpenAI introduced Copyright Shield in November 2023, covering ChatGPT Enterprise and the API and explicitly not the free and Plus tiers. Its service terms condition the cover on your behaviour. It does not apply where you disabled or ignored citation, filtering or safety features. Nor does it apply where you lacked rights to the input or fine-tuning files, or where the output was modified or combined with non-OpenAI products.
Google split its commitment in two in October 2023. A training data indemnity covers claims that Google's use of training data infringes third-party IP. A generated output indemnity covers claims against the output, and applies only if you did not intentionally create infringing output and use tools such as citation responsibly.
My reading is that these converge more than vendors imply. Each excludes customer-supplied inputs, each excludes outputs you modified, and each excludes deliberate infringement. If your risk is a model reproducing training data, all three cover you. If your risk is that your own data was not yours to use, none of them do.
The cap behind the indemnity
An indemnity is only as good as the liability cap behind it. Anthropic publishes both, which makes it the easiest of the three to evaluate without a sales call. Liability is limited to fees paid in the previous 12 months, indirect damages are excluded, and the cap does not apply to the Section K indemnity.
That last clause is the one to check in every AI contract. A capped indemnity and an uncapped one are different products wearing the same word. If the indemnity sits inside a 12-month fee cap and you spend 40,000 dollars a year, the defence you were sold is worth 40,000 dollars.
This is the clause I would negotiate first, ahead of price. Pricing moves every quarter anyway, as the analysis of falling token prices and rising bills shows. A cap set at signature governs the whole term.
Where this comparison is weakest
Published terms are not the terms a large buyer gets. Enterprise agreements are negotiated, and a real deal can carry an uptime commitment, a higher cap and a custom retention schedule none of these documents show. This is accurate for the standard commercial path, where most teams sit, and it understates what a seven-figure buyer can extract.
Two of the three vendors block automated retrieval. Anthropic's terms and documentation were read directly here. OpenAI's policy pages and Google Cloud's documentation return errors or render client side, so clauses from those two were verified through search extracts of the primary pages rather than by opening them. That is weaker than this publication's usual standard, and it is disclosed rather than hidden.
Uptime figures are vendors marking their own homework. Status pages reflect what each company counts as an incident, across different component sets. Useful for order of magnitude, useless for a credit calculation.
The competing view deserves a hearing. A buyer could argue that all four terms are second-order, that capability and cost dominate, and that switching is cheap enough to make contract risk manageable. That holds until an agent is embedded in a workflow, which is the point made in the piece on building against buying coding agents.
Frequently asked questions
Does OpenAI or Anthropic train on API data?
Neither does by default on paid commercial terms. OpenAI states that data sent to the API has not been used to train or improve its models since 1 March 2023 unless a customer explicitly opts in. Anthropic's commercial terms state that Anthropic may not train models on customer content from the Services, and that customers retain rights to inputs and own outputs. Free consumer tiers follow different policies.
How long does each AI provider keep my prompts?
Anthropic automatically deletes API inputs and outputs within 30 days, with exceptions for flagged content, legal requirements and features with customer-controlled retention. OpenAI retains API inputs and outputs for up to 30 days for abuse monitoring. Google logs paid Gemini API prompts and responses for a limited period to detect prohibited use, and stores Search grounding prompts for 30 days.
Which AI provider offers the best indemnity for copyright claims?
All three defend business customers against third-party IP claims on outputs, and the practical differences are small. Anthropic lists six carve-outs including patent practice and trademark use. OpenAI conditions its cover on using its citation, filtering and safety features. Google splits cover into a training data indemnity and a generated output indemnity. None covers inputs you had no right to use.
Do AI model providers offer an uptime SLA?
Not in their standard published terms. Anthropic's commercial terms contain no uptime percentage and no service credit. Availability commitments from the model labs are negotiated into enterprise agreements, so they vary by buyer. Clouds reselling the same models publish uniform credit-backed SLAs. Amazon Bedrock pays credits of 10%, 25% or 100% depending on the monthly uptime band achieved.
Is zero data retention available on all endpoints?
No, and this is the most common misunderstanding. Anthropic enables it per organisation and excludes batch processing, code execution, managed agents and its Covered Models. OpenAI enables it per endpoint for pre-approved customers, excluding the assistants, threads and conversations endpoints. Google requires disabling data caching and opting out of abuse-monitoring logging, and some advanced features remain outside it.
Should I buy models direct or through a cloud provider?
Buy direct if you want the newest models first and simpler commercial terms. Buy through a cloud if you need a published uptime remedy, an existing data processing agreement and a single procurement path. Note that on Amazon Bedrock and Google Cloud the cloud provider is the data processor, not the model lab, so the retention and compliance documentation that applies is the cloud's.
Where to start this week
Two things, both short.
First, find out which billing account every model key in your codebase belongs to, and confirm none is an unpaid tier. That takes 20 minutes and closes the only gap here where a provider is contractually allowed to read and train on your content.
Second, open your current agreement and locate two clauses: the liability cap, and whether the indemnity sits inside or outside it. If the indemnity is capped at 12 months of fees, write that number down next to your renewal decision, alongside the questions raised in the review of what the agent orchestration tools actually automate.
One thing to take to your next renewal
Ask each provider for the uptime remedy, the retention exclusion list and the liability cap in writing, before discussing price. A vendor who will put all three in an email has done the work.
References
- Anthropic, Commercial Terms of Service, effective 17 June 2025. Used for the training clause, output ownership, Section K indemnity carve-outs and the Section L.3 liability cap.
- Anthropic, API and data retention, retrieved 21 August 2026. Used for zero data retention scope, feature eligibility, Covered Models and session transcript retention.
- Anthropic Privacy Center, How long do you store my organization's data, retrieved 21 August 2026. Used for the 30-day deletion default and the flagged content periods.
- OpenAI, Data controls in the OpenAI platform, retrieved 21 August 2026. Used for the training position, 30-day abuse monitoring and zero-retention endpoint eligibility.
- Google, Gemini API Additional Terms of Service, retrieved 21 August 2026. Used for the paid and unpaid data-use split, human review and the 30-day grounding retention.
- Google Cloud, Protecting customers with generative AI indemnification, 13 October 2023. Used for the two-part indemnity and its conditions.
- Amazon Web Services, Amazon Bedrock Service Level Agreement, effective 4 October 2023. Used for the credit bands and the claim procedure.
- Engadget, OpenAI no longer has to preserve all of its ChatGPT data, with some exceptions, 23 October 2025. Used for the narrowing of the preservation order.
The weakest thing about this source base: OpenAI's policy pages and Google Cloud's documentation blocked automated retrieval. Clauses from those two were verified through search extracts of the primary pages rather than by reading them directly. Anthropic's documents were read in full. Terms change without notice.
Related reading