From Ritu Raj | Product & Market Analysis
The Copyright Cases That Could Reprice Training Data, and What They Cost Now
On this page
One AI copyright lawsuit has produced a number you can build a model on. Anthropic agreed to pay $1.5 billion to settle a class action brought by authors, and a federal judge granted final approval on 20 July 2026. That works out at roughly $3,000 for each claimed work. Every other live case is still arguing about whether the price should be $750 or $150,000.
Key takeaways
- The market now has one observed price for a copied book: about $3,000. The court itself noted that figure is four times the $750 statutory minimum for ordinary infringement, and far above anything proposed in the Google Books settlement.
- The legal line so far runs through acquisition, not training. Courts have been more willing to treat training as transformative than to forgive building a corpus out of pirated files. Provenance records are the asset, not the model weights.
- No appellate court has ruled on fair use in AI training. The Third Circuit heard argument in Thomson Reuters v. Ross on 11 June 2026 and has not decided. Until it does, every settlement is priced against a guess.
- Licensing is still cheaper than litigation by an order of magnitude. The largest reported publisher deal, News Corp with OpenAI, is worth over $250 million across five years. One settlement cost six times that in a single payment.
What the AI copyright cases are actually deciding
Almost every headline on this topic asks whether training on copyrighted work is legal. That is the wrong unit of analysis. The courts are not ruling on training in the abstract.
They are ruling on four separate questions, and the answers are diverging. Whether the copy made during acquisition was lawful. Whether the training use was transformative. Whether the output substitutes for the original. And whether a licensing market exists that the defendant harmed by not paying into it.
A company can lose on the first and win on the second. Anthropic did exactly that in 2025, and the settlement that followed was priced on the loss, not the win.
The direct-answer version
No court has issued a binding final answer on whether AI training is fair use. One district court held that training on lawfully acquired books was transformative while pirating those books to build a library was not. A second held that training a legal research tool on Westlaw headnotes was not fair use at all. Both are on appeal or settled.
Why the count of cases is not the number to watch
Public trackers put the number of US AI copyright suits somewhere between roughly 90 and 200, depending on whether they count every complaint naming an AI product or only training-data theories. The Chat GPT Is Eating the World tracker listed 138 US suits on 29 August 2026. Treat that spread as a warning about methodology rather than a measure of risk.
Volume tells you the topic is contested. It does not tell you what a copy of a book costs. Only two things do that: a judgment, or a settlement large enough to anchor the next negotiation. There is one of each so far.
The first case that put a price on a training corpus
Bartz v. Anthropic is the only AI copyright matter that has generated a number a finance team can use. Judge Araceli Martinez-Olguin granted final approval on 20 July 2026 in the Northern District of California.
The headline is $1.5 billion. The useful figure is the one underneath it.
The arithmetic that anchors every future negotiation
Authors receive roughly $3,000 for each eligible claimed work. Divide the fund by that figure and you get an implied corpus of about 500,000 works. That is the unit price of a pirated book, established by a court order rather than by a vendor claim.
Set it against the statutory range. Ordinary infringement carries a $750 minimum. Willful infringement runs to $150,000 per work. The settlement landed at four times the floor and about 2% of the ceiling.
My read is that the ceiling was never realistic and both sides knew it. A 500,000-work class at $150,000 each is $75 billion, which is larger than most of the companies being sued. What the settlement really priced was the probability of losing, not the damage.
What the settlement did not buy
Read the release carefully before treating this as closure. It covers Anthropic's acquisition and copying of works through 25 August 2025. Claims about model outputs and all future conduct survive untouched.
Anthropic also has to destroy the original files sourced from Library Genesis and the Pirate Library Mirror, along with copies derived from them. That is a data deletion obligation attached to a copyright settlement, and it is a template other plaintiffs will now ask for.
The settlement creates no binding precedent either. It resolves one class. The next plaintiff starts from zero on the law and from $3,000 a work on the economics, which is the reverse of what most coverage implied.
The split that matters is acquisition, not training
If you read only one distinction from 2025 and 2026, make it this one. Courts have been considerably more receptive to the argument that training is transformative than to the argument that any means of obtaining the corpus is acceptable.
That maps to a practical rule. Your exposure tracks your procurement records, not your model architecture. A lab that bought and scanned books sits in a different position from one that torrented them, even where the resulting weights are identical.
| Case | Court and status | What a ruling would set a price on |
|---|---|---|
| Bartz v. Anthropic | N.D. Cal. Settled, final approval 20 July 2026 | Already priced. About $3,000 per pirated book |
| Thomson Reuters v. Ross Intelligence | Third Circuit. Argued 11 June 2026, undecided | Whether training a competing tool on licensed data is fair use |
| In re OpenAI Copyright Litigation | S.D.N.Y. MDL, discovery through 2026 | News archives, and whether memorised output substitutes for the original |
| Kadrey v. Meta | N.D. Cal. Partly dismissed June 2025, seeding claims live | Distribution during torrenting, separate from training |
| Disney and others v. Midjourney | C.D. Cal. Discovery | Character output, with statutory damages sought up to $150,000 per work |
| Getty Images v. Stability AI | England and Wales High Court. Decided 4 Nov 2025 | Almost nothing. Getty dropped its main claims before closing |
Status as at 30 August 2026, compiled from the Norton Rose Fulbright 2026 case update, the Authors Guild, Courts and Tribunals Judiciary and contemporaneous reporting. Procedural posture in active litigation changes weekly.
Getty is the case people cite wrongly
Getty Images v. Stability AI is regularly described as a test of AI training in the United Kingdom. It was not, in the end. Getty abandoned its primary copyright infringement and database right claims before closing submissions.
Mrs Justice Joanna Smith rejected the remaining secondary copyright claim in her 4 November 2025 judgment. What Getty won was a trade mark point the court described as historic and extremely limited. Anyone citing the case as settling the training question has not read it.
The case that has not landed, and why it is the expensive one
The consolidated OpenAI litigation in the Southern District of New York is the matter with the widest commercial consequences. Judge Sidney Stein denied most of the motion to dismiss in March 2025 and declined to accept that training is inherently transformative.
Discovery has been unusually aggressive since. The court granted a motion to compel in March 2026, and the plaintiffs pushed successfully for access to a very large volume of ChatGPT conversation logs. That matters beyond this case: it establishes that output logs are discoverable evidence, which changes what every model provider should assume about retention.
No merits ruling had issued as of late August 2026. The honest position is that nobody outside the chambers knows the timing, and dates published in vendor trackers are estimates.
What makes this the expensive case is not the damages theory. It is that a finding of memorisation-driven substitution would apply to news archives, technical documentation and reference works at once, which is exactly the corpus enterprise models are most valuable for. The cost pass-through would land in inference and licensing lines that are already the weakest part of AI gross margin.
What licensing actually costs, when anyone will say
The licensing market exists, and it is much smaller than the litigation. That gap is the whole commercial story.
The largest reported publisher agreement is News Corp with OpenAI, valued at more than $250 million over five years according to Wall Street Journal reporting. Reddit disclosed a $60 million a year arrangement with Google in 2024, with a further agreement reported at around $70 million with OpenAI. Universal Music settled with Udio in October 2025 and converted the dispute into a licence, with terms undisclosed.
Compare those to one settlement of $1.5 billion. A five-year exclusive over the entire News Corp archive costs a sixth of what a single class action cost. Either the litigation is overpriced or the licences are, and my view is that the licences are underpriced because publishers negotiated before anyone knew what a court would award.
Where this argument is weakest
This post rests on a settlement, an unrendered appeal and a set of reported deal values. Each of those is softer than it looks.
One settlement is not a market price
Bartz priced a specific fact pattern: books obtained from known piracy libraries, by a well-capitalised defendant, in a class the court could define cleanly. Almost none of that transfers to images, music, code or news archives, where the works are harder to enumerate and ownership is often split across parties.
Anyone using $3,000 a work as a general planning figure is extrapolating from a sample of one. I am doing it in this post because it is the only observation available, and that is a weakness of the evidence rather than a strength of the analysis.
The licensing figures are reported, not disclosed
Not one of the publisher deal values above appears in a filing. They come from press reporting on private contracts. Reddit's arrangements are the closest to verifiable because the company is listed and discusses licensing revenue publicly.
The undisclosed deals could easily be larger than the reported ones. Selection bias runs one way here: the deals that leak tend to be the ones a party wanted known.
The case against writing any of this down yet
There is a respectable position that all of this is premature. No appellate court has ruled. The Third Circuit could hold that training a competing product on licensed data is fair use, which would cut the value of every publisher licence signed to date.
If that happens, the $3,000 figure becomes a historical curiosity and the licensing market repriced downward, not upward. I think that outcome is less likely than the alternative, given how the panel questioned both sides on market harm. I would not stake a procurement decision on my read of an oral argument, and neither should you.
What this changes if you buy AI, rather than build it
You are not the defendant in any of these cases. The copying happened upstream. Your exposure is contractual, and it is entirely within your control to price it.
Three things follow. First, the indemnity in your model contract is now the most valuable clause in it, and most buyers have never checked whether it covers training data claims or only outputs. That is the same category of gap covered in the review of how the major model providers' terms actually differ, and it belongs on the list of contract clauses a finance team should be reading before renewal.
Second, provenance is becoming a disclosure obligation rather than a courtesy. Article 53 of the EU AI Act requires providers of general-purpose models to publish a summary of training content using a template the AI Office issued in July 2025. It applied to new models from 2 August 2025, with existing models given until 2 August 2027. That is a hard date, and it sits alongside the wider transparency obligations the Act imposes on deployers.
Third, the same logic applies inward. If your own product trains on customer data, the consent question you are creating for yourself is the one covered in the analysis of what note-taking tools actually claim over recorded conversations.
None of this is a reason to slow down adoption. It is a reason to write two paragraphs into your vendor review that were not there last year, and the cost of doing so is an hour.
Frequently asked questions
Is training AI on copyrighted work fair use?
No court has issued a final answer that binds anyone. A California judge held in 2025 that training on lawfully bought books was transformative, while holding that building a library from pirated files was not. Thomson Reuters v. Ross went the other way on a non generative research tool. The Third Circuit heard argument on that case in June 2026 and has not yet ruled.
How much did Anthropic pay to settle the authors' copyright case?
Anthropic agreed to pay 1.5 billion dollars to settle Bartz v. Anthropic, and Judge Araceli Martinez-Olguin granted final approval on 20 July 2026. The fund works out at roughly 3,000 dollars for each eligible claimed work. The court cut the requested attorney fees to about 101.6 million dollars and trimmed each class representative service award from 50,000 to 15,000 dollars.
What is the status of the New York Times case against OpenAI?
The Times case sits inside a consolidated multidistrict proceeding in the Southern District of New York before Judge Sidney Stein. He denied most of the motion to dismiss in March 2025 and rejected the argument that training is inherently transformative. Discovery has run through 2026, including an order to produce ChatGPT conversation logs. No merits ruling had issued as of late August 2026.
Did Getty Images win its case against Stability AI?
Mostly no. Getty dropped its primary copyright and database right claims before closing submissions, and Mrs Justice Joanna Smith rejected the secondary copyright claim in her 4 November 2025 judgment. Getty won a narrow trade mark point the court called historic and extremely limited. The judgment therefore says very little about whether training on copyrighted work is lawful in the United Kingdom.
How much does it cost to licence content for AI training?
Prices are mostly private, so the public record is thin. The largest reported publisher deal is News Corp with OpenAI, valued at more than 250 million dollars over five years according to the Wall Street Journal. Reddit disclosed a 60 million dollar a year agreement with Google in 2024. Most other agreements, including the Universal Music settlement with Udio, carry undisclosed terms.
Am I liable if my company uses an AI model trained on copyrighted data?
Direct exposure sits with whoever copied the work, which is usually the model developer, not you. The practical risk for a buyer is contractual rather than statutory. Read the indemnity in your model contract, check whether it covers training data claims or only output claims, and check the cap. If the answer is not written down, assume the risk sits with you and price it accordingly.
Where to start this week
Pull the two contracts that matter most and read one clause in each. In your largest model provider agreement, find the indemnity and write down whether it names training data claims, output claims, or neither. In your largest data processing agreement, find whether your own customer content can be used for training and under what notice.
Then set a single calendar entry for the Third Circuit decision in Thomson Reuters v. Ross. It is the first appellate word on fair use in AI training, and it will move licensing prices in one direction or the other within a quarter of landing. Everything else on this page is a footnote to that ruling.
Related on the money side
Licensing costs land in the same line as compute. See how the underlying spending is justified in the payback math on 2026 AI capex, and what it does to margins in the breakdown of inference costs.
References
- Authors Guild, Court grants final approval of $1.5 billion Anthropic copyright settlement, July 2026. Used for the settlement amount, approval date, judge, per-work figure, fees and service awards.
- Publishing Perspectives, Court grants final approval to landmark $1.5 billion Anthropic settlement, July 2026. Used for the per-work payout and the comparison to the statutory minimum.
- Norton Rose Fulbright, An update on AI copyright cases in 2026, 2026. Used for the procedural status of Ross, Kadrey, the OpenAI MDL and Disney v. Midjourney.
- Courts and Tribunals Judiciary, Getty Images v. Stability AI [2025] EWHC 2863 (Ch), 4 November 2025. Used for the citation, date and judge.
- LawSites, At 3rd Circuit, judges press ROSS and Thomson Reuters on fair use, June 2026. Used for the oral argument date and the panel's focus on market harm.
- Nieman Lab, OpenAI and News Corp strike a content deal valued at over $250 million. Used for the News Corp licensing figure, which originates in Wall Street Journal reporting.
- European Union, AI Act Article 53, obligations for providers of general-purpose AI models. Used for the training content summary requirement and the 2 August 2027 transitional date.
- Music Business Worldwide, Universal Music settles Udio lawsuit, strikes deal for licensed AI music platform, October 2025. Used for the settlement-to-licence conversion and the absence of disclosed terms.
The weakest part of this source base is the licensing figures. None appear in a filed contract or a regulatory disclosure, so every deal value here is press reporting on private terms. Procedural status is current as at 30 August 2026 and changes weekly in active litigation.
Related reading