From Aryan Vatsa | Product & Market Analysis
AI Data Retention Policies, Ranked: What Ten Vendors Guarantee on Delete
On this page
Amazon Bedrock scores highest on AI data retention, and the two frontier labs do not come second. The reason is a change most buyers have not read. Since 9 June 2026 Anthropic retains prompts and outputs from its covered models for 30 days on every platform. That policy overrides zero data retention agreements customers already signed. A retention promise is now a per-model fact with an expiry date.
Key takeaways
- Zero data retention is no longer a property of your contract. Anthropic's covered models policy retains prompts and outputs for 30 days and states that it overrides existing zero retention setups in Claude Console, on AWS Bedrock, on Google Cloud and in Microsoft Foundry.
- The cloud resellers publish better controls than the labs they resell. Bedrock exposes retention as an API value you can read back and lock with an IAM policy. Azure exposes a ContentLogging flag you can check yourself. Neither lab offers an equivalent.
- Google's paid Gemini API ranks lowest of the enterprise tiers on disclosure. Its terms say prompts are logged for a limited period of time and never name the period, while the consumer Gemini app publishes 18 months, 72 hours and 3 years.
- Consumer terms are a different product, not a lighter version of the same one. Claude Free, Pro and Max move retention from 30 days to 5 years if you leave training on. Claude for Work and the API are excluded from that change entirely.
Which AI vendor has the strongest data retention policy
Amazon Bedrock ranks first on the rubric below, scoring 9 out of 10. It defaults to zero retention, exposes the setting as an API value you can read back, and lets you freeze it organisation wide with a service control policy. Microsoft Foundry and the OpenAI API tie on 8. Meta AI scores 0.
This piece is written for the head of delivery or operations lead who has to fill in the security questionnaire and then actually configure the thing. That person carries the gap between what a policy page says and what a tenant is set to.
Why the resellers beat the labs
The labs write the retention policy. The resellers have to prove theirs to auditors who do not accept a marketing page as evidence. That difference in obligation shows up directly in the products.
My position is that verifiability now matters more than the number itself. A 30-day window you can read from an API beats a zero-day window you can only read from a PDF.
The rubric, and what it deliberately does not score
Ten vendors, five criteria, 0 to 2 points each, 10 points maximum. Every score below traces to a vendor document read on 1 September 2026, listed in the references. The rubric is mine, the underlying facts are theirs.
| Criterion | What earns 2 points | What earns 0 |
|---|---|---|
| Published window | A specific retention period stated in days on the paid tier | No period stated anywhere in the terms |
| Training default | Your content is excluded from model training with no action required | Training is on unless you find the switch |
| Zero retention path | Available by configuration, not by sales approval | No zero retention concept exists |
| Durability | The vendor cannot unilaterally override the arrangement | A published policy already overrides it |
| Verifiability | You can read the live setting yourself through a console or API | You can only read a policy page |
Excluded on purpose: price, model quality, certifications, breach history and negotiated enterprise addenda. Certifications tell you a process exists. They do not tell you how long your prompts sit on disk.
Disclosure carries the same weight as the training default. That is a deliberate choice and it moves the ranking. A vendor that trains on your data and says so is easier to govern than a vendor that does not train and will not say how long it keeps anything.
Every major AI vendor's data retention policy, ranked
Read the score as a governance score, not a safety score. It measures how much of the answer you can obtain and check without a sales call.
| Rank and vendor | Default retention | Trains on your content | Zero retention | Score |
|---|---|---|---|---|
| 1. Amazon Bedrock | None by default. Up to 30 days for models needing review | No, and content is not shared with the model provider | Default mode, set by API | 9 |
| 2. Microsoft Foundry | Flagged content only, no day count published | No, stated explicitly | Modified abuse monitoring, by application | 8 |
| 2. OpenAI API | Up to 30 days for abuse monitoring | No, unless you opt in | By approval, on listed endpoints only | 8 |
| 4. Cursor, Enterprise | Providers store nothing under Cursor's agreements | No | Privacy Mode on by default | 7 |
| 5. Anthropic API | 30 days, and 30 days for covered models regardless | No, under commercial terms | By approval, now overridden for covered models | 6 |
| 5. GitHub Copilot Business | Nothing in the IDE. 28 days for chat messages | No | Via GitHub's agreements, not yours | 6 |
| 5. Microsoft 365 Copilot | Whatever your Purview policy says | No, stated explicitly | Not offered as a concept | 6 |
| 8. Claude Free, Pro, Max | 30 days, or 5 years if training is left on | Yes, unless you opt out | Not offered | 4 |
| 9. Google Gemini API, paid | A limited period of time, not quantified | No on paid, yes on unpaid | Not offered | 3 |
| 10. Meta AI | Not published | Yes, and chats feed ad personalisation | Not offered | 0 |
Rows describe published defaults for the tier named, not negotiated enterprise addenda. A large customer can buy terms that differ from every row here, which is precisely why the published default is the useful benchmark.
The override nobody priced into their contract
Until this year, zero data retention behaved like a contract term. You asked, you were approved, and the answer held. That is no longer true for the newest models.
Anthropic's covered models policy took effect on 9 June 2026. It states that prompts submitted to and outputs generated by covered models are retained for 30 days to support safety work, on every platform where those models are offered. The policy names Claude Fable 5, Fable 5.1, Mythos 5 and Mythos 5.1.
The reach is the part to read twice. It applies to organisations that already had zero data retention in Claude Console and in Claude Code under a Claude Enterprise agreement. It also covers those reaching Claude through AWS Bedrock, Google Cloud or Microsoft Foundry. Buying through a hyperscaler does not route around it.
Amazon documents the same requirement from the other side, and more bluntly. Bedrock's retention modes run from none through default to aws_review. Claude Fable 5 and Fable 5.1 require aws_review, so an account set to none sees the model as unavailable and its requests are blocked.
That is a clean statement of the new trade. You can have the frontier model or you can have zero retention. Choosing both is not on the menu unless you qualify for an exemption.
The exemptions have dates on them
Both AWS and GitHub describe the same escape hatch and both put a clock on it. AWS says customers eligible for the Enterprise Frontier Safeguards programme receive zero retention through 31 December 2026, after which all traffic is retained for up to 30 days. GitHub says eligible enterprises can request zero retention for the same models through the end of 2026 under a time-bound exemption.
If your data protection impact assessment rests on a zero retention arrangement, check whether it expires this December. I would put that date in a calendar today rather than discover it in a renewal.
Consumer terms and enterprise terms are different products
The single most common procurement error I see is treating a vendor as one policy. Inside the same brand, the consumer tier and the commercial tier can differ by two orders of magnitude on retention.
Anthropic is the clearest example because it published both numbers. Its August 2025 consumer update moved Free, Pro and Max users from the existing 30-day period to five years where the user allows training. Existing users were required to choose by 8 October 2025. The same announcement excludes Claude for Work, Claude for Government, Claude for Education and API use, including through Amazon Bedrock and Google Cloud.
| Vendor | Consumer tier | Commercial tier |
|---|---|---|
| Anthropic | 5 years in the training pipeline if training is on, 30 days if off | Deleted within 30 days, excluded from the consumer training change |
| 18 months default, 72 hours with activity off, up to 3 years if human reviewed | Paid API not used to improve products, logged for an unnamed period | |
| Microsoft | Not covered by the commercial commitments | Prompts, responses and Graph data not used to train foundation models |
| Meta | AI conversations personalise content and ads from 16 December 2025 | No comparable enterprise assistant tier on the same terms |
The free tier is where training rights actually diverge
Google's Gemini API terms draw the line explicitly. On unpaid services, content submitted and generated is used to provide, improve and develop Google products, and human reviewers may read, annotate and process it. On paid services, prompts and responses are not used to improve products.
Read that as a pricing decision rather than a privacy decision. The free tier is paid for in data, which is a fair trade as long as nobody in your team believes they are on the paid terms.
Meta's own announcement confirms that interactions with its AI features personalise content and ads from 16 December 2025. It adds that sensitive topics such as religious views, sexual orientation, political views and health are not used for ad targeting. What the announcement offers is ad preference controls, not a switch that stops AI conversations feeding recommendations.
Press reporting says the change was not rolled out in the EU, the UK and South Korea. Meta's post says only that it is rolling out in most regions and names no exclusions, so treat the regional list as reported rather than confirmed. That gap between what is reported and what is stated is itself a finding, and it is the reason Meta scores 0 on disclosure.
What deletion does not delete
Every vendor here lets you delete a conversation. None of them means by that what a user assumes it means. Anthropic publishes the most detailed version of the truth, so its numbers make the best worked example.
A deleted Claude conversation is removed from your chat history immediately and from back-end storage within 30 days. If the conversation triggered a usage policy violation, inputs and outputs are retained for up to 2 years, and the trust and safety classification scores derived from them for up to 7 years.
So a single deletion clears four different things on four different clocks. That is not a criticism of Anthropic. It is the shape of the answer at every vendor, and Anthropic is unusual in printing it.
The endpoint you use changes the answer
OpenAI publishes the most operationally useful version of this. Its zero retention arrangement covers chat completions, responses, embeddings, images, audio and moderations. It does not cover conversations, assistants, threads, vector stores, files, batches, evals, fine tuning jobs or videos.
An engineer who moves a feature from chat completions to the conversations endpoint has quietly changed the retention posture of that feature. Nothing in the code review will say so. This is the same failure mode covered in the comparison of the four provider contract terms that decide your risk, seen from the implementation side rather than the contract side.
Where this ranking is weakest
Four objections, all of them fair.
The rubric rewards disclosure more than most buyers would
Google's paid Gemini API does not train on your content, which is the outcome most buyers actually want, and it still ranks ninth. Reweight the published window criterion downward and it rises above the Claude consumer tier immediately. I keep the weight because an unnamed window cannot be tested, and an untestable control is not a control. Reasonable people will score this differently.
Published defaults are not what large buyers sign
Every vendor here will negotiate. A company spending eight figures gets terms no policy page describes, which means this table is least accurate for exactly the buyers with the most at stake. The table is a floor and a starting position, not a prediction of your final contract. What to negotiate from here is set out in the piece on the data processing agreement clauses worth arguing over.
Policy pages change without notice
Anthropic's covered models policy did not exist in early 2026. Meta's ad personalisation change took effect in December 2025. Every score here is a reading on 1 September 2026 and several will be wrong within two quarters. Anything built on this table needs a re-verification date attached.
One vendor is scored on an absence
Microsoft Foundry scores 0 on the published window criterion. The abuse monitoring and data privacy pages I read do not state a retention period in days for flagged content. That may be a documentation gap rather than a policy gap, and Microsoft may state the figure in the Product Terms or the Data Protection Addendum. I did not find it in the public documentation, and I scored what is public.
Six questions to put in your vendor questionnaire
Most AI security questionnaires ask whether the vendor trains on customer data. That question stopped discriminating between vendors around the time everyone learned to answer no. These six still separate them.
| Ask this | What a good answer contains | What a weak answer looks like |
|---|---|---|
| How many days do you retain prompts and outputs on our tier? | A number, and the storage system it applies to | A limited period, appropriate period, or as required |
| Which of our endpoints or features are excluded from that? | A named list, like OpenAI's eligible endpoint list | An assurance that it applies to the service |
| Can you change our retention terms without our consent? | The circumstances, named, with the notice period | Silence, or a reference to the general terms |
| How do we verify the current setting ourselves? | A console path or API call we can run today | A support ticket or an account manager |
| What survives a deletion request, and for how long? | Separate answers for logs, flags and derived scores | One number covering everything |
| Does a new model change any of the above? | A stated policy for new model launches | Confidence that nothing will change |
The sixth question is new this year and it is the one I would add first. Before June 2026 it would have looked paranoid. It is now the question that would have caught the covered models change before it reached a live workload. The same logic applies further down the stack, which is why the agent procurement checklist treats model substitution as a contract event rather than an engineering detail.
Where the questionnaire will not help you
None of this reaches the personal accounts your team already uses. A consumer Claude or Gemini login inside a work browser sits outside every control described here, on the retention terms in the consumer rows of the table above. That exposure is quantified in the piece on what unsanctioned AI use costs when it breaches, and the boundary problem is mapped in the analysis of personal AI accounts outside the perimeter.
Meeting recorders deserve their own pass, because consent and retention interact there in a way they do not elsewhere. That case is handled separately in the review of note taker consent and training rights.
Frequently asked questions
How long does OpenAI keep API data?
OpenAI retains API inputs and outputs for up to 30 days for abuse monitoring, after which they are removed unless retention is legally required. Assistants API data such as threads and vector stores is retained for 30 days after you delete it through the API, and conversations endpoints retain data until deleted. Data sent to the API is not used to train OpenAI models unless you explicitly opt in.
Does Anthropic train on my data?
It depends on the product. Under commercial terms, covering Claude for Work and the Anthropic API, your content is not used for training. On the consumer Free, Pro and Max plans, chats and coding sessions are used to train models unless you opt out, and choosing to allow training extends retention from 30 days to five years. Existing consumer users had to make that choice by 8 October 2025.
What is zero data retention in AI?
Zero data retention means the provider does not store your inputs or outputs after generating a response. In practice it is narrower than it sounds. It usually applies only to named endpoints, often requires sales approval, and can be overridden. Anthropic's covered models policy retains prompts and outputs for 30 days on every platform from 9 June 2026, including for organisations that already had zero retention arrangements.
Which AI vendor has the best data retention policy?
On the five-criterion rubric in this post, Amazon Bedrock scores highest at 9 out of 10. It defaults to zero retention, lets you set the mode through an API, and lets you enforce it organisation wide with a service control policy. Microsoft Foundry and the OpenAI API tie on 8. The ranking measures published and verifiable controls, not model quality or security posture generally.
Can I opt out of AI training on my data?
On paid enterprise tiers you usually do not need to, because training is off by default at OpenAI, Anthropic, Microsoft and on Google's paid API. On consumer tiers it varies. Claude Free, Pro and Max offer an opt out in privacy settings. Google's unpaid Gemini API content is used to improve Google products and may be read by human reviewers. Meta AI offers ad preference controls rather than an opt out of AI chat personalisation.
Does deleting a ChatGPT or Claude conversation actually delete it?
Not immediately and not entirely. Anthropic removes a deleted conversation from your chat history at once and from back-end storage within 30 days. If the conversation was flagged for a usage policy violation, inputs and outputs are kept for up to 2 years and the resulting safety classification scores for up to 7 years. Google keeps human-reviewed Gemini chats for up to three years even after you delete your activity.
Where to start this week
Pick your two largest AI vendors and answer question one from the table above for each of them. Do it in writing, with the URL of the page you read it on. If you cannot find a number, that is the finding, and it belongs in your risk register rather than in a follow-up email.
Then check one thing in your own tenant. If you are on Bedrock, read back your retention mode. If you are on Azure, check whether ContentLogging is false. Configuration drifts away from policy quietly, and the gap between them is the only part of this you control directly.
Keep this one current
Every score here is a reading on 1 September 2026 and the covered models change shows how fast these move. Re-verify before any renewal, and start with the four provider contract terms that decide your risk.
References
- Anthropic Privacy Center, Data retention practices for Covered Models, effective 9 June 2026, and How long do you store my organization's data, read 1 September 2026. Used for the covered models retention period, its override of zero retention arrangements, and the 30 day, 2 year and 7 year figures.
- Anthropic, Updates to Consumer Terms and Privacy Policy, 28 August 2025. Used for the five year consumer retention figure and the excluded products.
- OpenAI, Your data, read 1 September 2026. Used for the 30 day API window and the eligible and ineligible endpoint lists.
- Amazon Web Services, Amazon Bedrock data retention and Amazon Bedrock abuse detection, both read 1 September 2026. Used for the retention modes, the model availability rule, the service control policy example, the zero operator access model and the Enterprise Frontier Safeguards end date.
- Google, Gemini API Additional Terms of Service and the Gemini Apps privacy hub, both read 1 September 2026. Used for the paid and unpaid data usage split, the unquantified logging period, and the 18 month, 72 hour and 3 year figures.
- Microsoft, Data, privacy and security for Microsoft Copilot, updated 18 August 2026, and Data, privacy and security for Foundry models sold by Azure, updated 5 June 2026. Used for the training commitments, the Purview retention path and the ContentLogging verification method.
- GitHub, Hosting of models for GitHub Copilot, read 1 September 2026. Used for the zero retention agreements and the Claude Fable exception.
- Cursor, Privacy and data governance, read 1 September 2026, and Meta, Improving your recommendations on our apps with AI at Meta, October 2025. Used for the Enterprise Privacy Mode default, the cloud agent storage behaviour, the 16 December 2025 date and the excluded sensitive topics.
Weakest thing about this source base: the scores are a reading of public documentation on a single day, and public documentation is not the same object as a signed enterprise agreement. Two claims are weaker than the rest. The regional exclusions for Meta's ad personalisation change come from press reporting rather than from Meta's own post. The absence of a published retention figure for Microsoft Foundry is an absence in the documentation read, not proof that no figure exists elsewhere.
Related reading