From Ritu Raj | Product & Market Analysis
AI Content Licensing Deals Tallied: Who Paid, and What the Terms Show
On this page
Only one AI content licensing deal has its price printed in a securities filing. Reddit told the SEC in February 2024 that it had signed data licensing arrangements worth $203.0 million in aggregate. We tallied 14 AI licensing deals with any public number attached, and only 5 rest on a figure the seller or buyer stated itself. The rest of the market price for training and retrieval rights is built from leaks.
Key takeaways
- Disclosed AI licensing terms cluster between $10 million and $66 million a year per seller. Reddit, News Corp, the New York Times, Dotdash Meredith, Wiley and Taylor & Francis all land in that band on an annualised basis.
- Most of the "market price" is reported, not disclosed. Of 14 deals with a public number, 8 rest only on press reports citing unnamed sources, and 1 has no figure at all.
- A licensed book costs more than a pirated one settled. HarperCollins titles reportedly license at $5,000 each, while the Anthropic settlement pays roughly $3,000 per work, four times the $750 statutory minimum.
- Retrieval rights now carry the price, not training rights alone. The 2025 and 2026 deals from Amazon, Meta and Perplexity pay for real-time display and summaries, which makes them recurring rather than one-off dataset sales.
The short answer: AI companies have paid individual publishers and platforms roughly $10 million to $66 million a year for content licences, based on the few terms that are public. Reddit's filing is the hardest number. News Corp, the New York Times and Dotdash Meredith fall inside that band. Per book, licences run near $5,000 per title.
This post is written for founders and directors who own a content or data asset and are deciding whether to license it, and at what price. That reader needs comparables, not commentary, so the ledger comes first.
What counts as a disclosed AI licensing deal
A licensing deal here means a contract where an AI developer pays for access to text, images or video. It may pay for training, for retrieval into live answers, or both. Training rights cover building a model. Retrieval rights cover showing or summarising current content inside a product like ChatGPT, Alexa or Meta AI.
Most of these contracts are confidential. A number becomes public in one of three ways, and they are not equally reliable. We graded every deal on which route its number took.
Grade A: the company, a filing or a court said it. The number appears in a securities filing, an earnings call, a results release or a court order. Reddit's S-1, IAC's disclosures on Dotdash Meredith, and Wiley's earnings calls sit here. These figures are still incomplete, because companies disclose aggregates and minimums, but someone is accountable for them.
Grade B: press reports citing people familiar. The number was reported by an established outlet, usually The Wall Street Journal, Bloomberg, Reuters or the Financial Times, citing unnamed sources. The companies declined to confirm. Grade B is citable. It is also how most of the market's price signal was formed, which is the uncomfortable part.
Grade C: announced with no number. The parties confirmed a deal exists and said nothing about money. That is the majority of AI publisher deals by count. We included one Grade C entry, the Meta news bundle, because it shows how the newest deals are structured. We excluded dozens of others that carry no number at all.
The ledger: 14 AI licensing deals and their terms
Where a figure is a minimum, a range or a contested reading, the row says so. Annualised values divide a reported total by the reported term, and that is our arithmetic, not the company's.
| Seller and buyer | Public terms | Rights | Grade |
|---|---|---|---|
| Reddit, unnamed buyers (Jan 2024) | $203.0M aggregate, 2 to 3 year terms, at least $66.4M in 2024 | Data access, training | A: S-1 filing. |
| Reddit and Google (Feb 2024) | About $60M a year | Data API, training | B: Reuters, one source. |
| News Corp and OpenAI (May 2024) | More than $250M over 5 years, includes OpenAI credits | Archive and current content | B: WSJ. |
| Axel Springer and OpenAI (Dec 2023) | "Tens of millions of euros", annual per FT, over 3 years per Bloomberg | Archive plus current summaries | B: conflicting. |
| Dotdash Meredith and OpenAI (May 2024) | At least about $16M a year fixed, plus a variable component | Training and display | A: IAC disclosures. |
| Taylor & Francis and Microsoft (2024) | About $10M | Part of library, training | B: Bloomberg Law. |
| Wiley and unnamed tech firms (FY2024 to FY2026) | $23M, $40M and $49M by fiscal year, one $18M agreement | Content for training and AI products | A: earnings calls. |
| HarperCollins and Microsoft (Nov 2024) | $5,000 per title, split 50/50 with the author, opt-in | Select nonfiction backlist, training | B: Bloomberg. |
| Shutterstock, multiple buyers (2023) | About $104M of 2023 data revenue | Images and metadata, training | B: analyst reports. |
| New York Times and Amazon (May 2025) | $20M to $25M a year | Training, summaries in Alexa | B: WSJ. |
| Perplexity publisher programme (Aug 2025) | $42.5M initial pool, 80% of Comet Plus revenue to publishers | Retrieval, citation, agent use | A: company. |
| Disney and OpenAI (Dec 2025) | $1B equity investment plus warrants, 3-year licence, fee not stated | 200+ characters in Sora | A: company, partial. |
| Meta and CNN, Fox News, USA Today, People Inc, others (Dec 2025) | Multiyear, terms not disclosed | Real-time news in Meta AI | C: none. |
| News Corp and Meta (Mar 2026) | Up to $50M a year, about 3 years | Retrieval and training | B: WSJ. |
Grade A means the figure came from the company, a filing or a court. Grade B means press reporting from unnamed sources. The Reddit and Google row and the Reddit aggregate row likely overlap, because the filing does not name its buyers. Two News Corp rows involve a WSJ report about its own parent company.
Read the Grade column before the money column. The two largest press figures, News Corp with OpenAI and News Corp with Meta, were both reported by The Wall Street Journal, which News Corp owns. That does not make them wrong. It does mean the seller's own newsroom set the public anchor price for news licensing, and you should weigh that.
AI publisher deals on an annual basis
Licensing terms are quoted in different shapes: totals over five years, annual minimums, per-title fees and revenue pools. To compare them, convert each to a single year where the reported terms allow it. Where they do not, leave the deal off the chart rather than guess.
The ceiling is set by platforms, not newspapers
The two highest bars belong to Reddit. A platform with a live, high-volume archive of human conversation commands at least as much as the largest English-language newspaper group.
News Corp's two deals sit just below. If both WSJ reports are accurate, News Corp alone draws up to roughly $100 million a year from two AI buyers.
The floor is about 1% of a publisher's revenue
Two deals at the bottom of the band have been sized against the seller. The Amazon payment to the New York Times is nearly 1% of the Times's 2024 revenue, per the WSJ. An analyst put the OpenAI payment to Dotdash Meredith at about 1% of that publisher's revenue, according to Engadget's report on the IAC disclosure.
My read: 1% of revenue is the number that matters most for a publisher deciding whether to sign. It is meaningful money and nowhere near a replacement for the search and social traffic AI answers may displace. Anyone selling a licence as a new business line should be shown that ratio first.
Training data licensing priced per book
Book deals are the only part of the market quoted per unit, which makes them the cleanest comparables for anyone pricing a catalogue.
The HarperCollins agreement with Microsoft pays $5,000 per title, split evenly with the author, according to Sherwood's report of the Bloomberg story. It covers selected nonfiction backlist titles, and authors who decline are left out. The Authors Guild objected that a 50/50 split gives the publisher too much for a training licence.
Why a licence costs more than a settlement
A licence buys things a settlement does not. It buys a clean provenance record, curated titles the buyer selected, and agreed output limits. Reporting on the HarperCollins deal said outputs are capped at 200 consecutive words or 5% of a book. A settlement buys only release from past claims.
The Anthropic settlement releases claims about past acquisition and copying through 25 August 2025 only. It does not cover output claims or future conduct. So the $3,000 figure is the price of a past liability, not a forward right to use anything.
Academic publishers price by catalogue, not by title
Academic publishers have not disclosed per-title rates. They report totals. Wiley reported $40 million of AI licensing revenue in fiscal 2025, up from $23 million, including a single $18 million agreement with a multinational tech customer. Publishers Lunch reported $49 million for fiscal 2026, with recurring AI revenue of $8 million inside that.
That recurring line is the number I would watch. In fiscal 2026 most of Wiley's AI revenue still came from one-off dataset sales. Bloomberg Law similarly described Informa's 2024 AI income as $75 million of non-recurring data access. A one-off sale is a windfall. It is not a business.
Reddit's licensing numbers, the only ones in a filing
Reddit is the reference point because it is the only seller that put contract value into a securities filing. Its S-1 said that in January 2024 it entered data licensing arrangements with an aggregate contract value of $203.0 million and terms of 2 to 3 years. It expected a minimum of $66.4 million of that to be recognised in 2024, as TechCrunch reported from the filing.
The filing did not name the buyers. Reuters reported, citing one source, that the Google contract was worth about $60 million a year. The S-1 risk factors added that "substantially all of the contract value associated with our licensing revenue is derived from one of our partners".
Reddit's 2025 Form 10-K reports total revenue of $2,202.5 million, up 69%. It says other revenues increased "as a result of content licensing agreements executed in 2024 and 2025". The same risk factor now reads "two of our partners" instead of one, consistent with the second reported buyer, OpenAI.
The $66.4 million 2024 minimum was about 5.1% of Reddit's $1,300.2 million in 2024 revenue. Advertising was $2.1 billion of 2025 revenue, per the auditor's report. Licensing is material to Reddit's story and small in its income statement.
The most useful sentence in the filing is not the $203 million. It is the admission that almost all contract value comes from one or two buyers. There are perhaps five or six buyers with the budget to pay eight figures, and they all know what the others paid.
How licensing prices compare with litigation outcomes
Licensing prices are negotiated in the shadow of the courts. If a buyer believes training is fair use, the rational licence price falls toward zero. If courts disagree, it rises toward statutory damages. Three outcomes now bound that range. Our breakdown of the copyright cases that could reprice training data covers the legal reasoning.
| Case | Outcome | What it does to licence prices |
|---|---|---|
| Bartz v. Anthropic, N.D. Cal. | $1.5B settlement, final approval 20 July 2026, about $3,000 per work | Sets a floor for pirated acquisition. Says nothing about lawful copies. |
| Thomson Reuters v. Ross, 3rd Cir. | Copying 2,243 Westlaw headnotes to train a legal AI held not fair use, affirmed 29 September 2026 | Raises prices for content used to build a competing product. |
| Getty Images v. Stability AI, UK High Court | Secondary infringement claim rejected 4 November 2025, model weights not an infringing copy | Lowers UK leverage for image libraries on model distribution. |
| Where licensing wins | Licences cover retrieval, display and future use, which no settlement has released | Recurring retrieval rights have no litigated substitute. |
Divide $1.5 billion by the roughly 482,000 eligible works and you get about $3,100 per work. Then compare it with the HarperCollins licence at $5,000. Settling after the fact cost Anthropic less per title than licensing up front costs Microsoft. The $1.5 billion figure frightens people because of its total, not its unit price.
I disagree with the common read that this settlement proves licensing is the cheap option. Per unit, it suggests the opposite. What it proves is that pirated acquisition is the expensive mistake, which is a narrower lesson.
The Third Circuit's Ross decision is the first federal appellate ruling on fair use in AI training, per IPWatchdog's report. Commentators note it distinguished generative systems from Ross, which returned existing passages. That narrowing matters, but it hands specialist publishers whose content feeds a competing tool a stronger hand than they had a month ago.
Where this tally is weakest
A ledger that claims to set a market price should state where it cannot.
Most rows are not verifiable from primary documents
8 of 14 rows are Grade B. We opened the reporting for each, but the underlying contracts are confidential. If the WSJ's sources overstated the News Corp figures, the top of the band moves down by half.
Survivorship shapes the sample
Deals leak when someone benefits from the number being public. A seller wants a high figure out. A buyer rarely wants any figure out. So the reported deals may skew high, and the dozens of undisclosed deals may sit well below the band. The opposite is possible too: OpenAI equity and compute credits, which News Corp's deal reportedly includes, may make headline cash look smaller than the real consideration.
The units do not match
Annualising a five-year total assumes even payments. Reddit's $66.4 million is a minimum across buyers. Wiley's figures are fiscal-year revenue across several buyers. Disney received a licence fee nobody disclosed alongside a $1 billion equity investment going the other way. Treat the annualised chart as a rough band, not a price list.
Pricing your own content licence from these comps
If you own a content or data asset, the ledger gives you three anchors and one warning. Our view on whether to let crawlers in at all sits in the robots.txt decision for AI crawlers, and that decision comes before pricing.
| Your asset | Closest comparable | What it implies |
|---|---|---|
| Live user-generated discussion | Reddit and Google | Annual fee, recurring, high concentration risk. |
| Current news and archive | News Corp, NYT, Dotdash Meredith | Roughly 1% of revenue, retrieval rights included. |
| Books or long-form backlist | HarperCollins and Microsoft | Per-title fee, opt-in, output caps. |
| Specialist or professional reference | Wiley, Taylor & Francis | Mostly one-off dataset sales, push for recurring terms. |
Sell retrieval, not just training
Training rights are a one-time event. Once a model has learned from your archive, the buyer has little reason to renew. Retrieval rights to current content renew because the content goes stale. The Amazon, Meta and Perplexity structures all pay for display and summaries, and that is the side of the contract I would negotiate hardest on.
Perplexity's model is the clearest example of recurring pricing. It committed a $42.5 million initial pool and 80% of Comet Plus subscription revenue to publishers, paid by traffic, citations and agent use. The pool is small. The per-use structure is the right shape.
Citation and traffic are a separate question from licence fees. How AI answers choose which sources to cite is covered in the playbook on Reddit and Perplexity citations, and a licence does not guarantee either.
The News Corp deal with OpenAI reportedly included credits for News Corp's use of OpenAI technology. Disney took the opposite position, investing $1 billion in OpenAI alongside its licence. When money flows both ways, the headline fee is not the price. The pattern is familiar from the circular deals across the AI supply chain, and it deserves the same scepticism.
Frequently asked questions
How much do AI companies pay publishers for content licensing?
Based on public terms, AI companies pay individual publishers roughly $10 million to $66 million a year. Reddit disclosed $203.0 million of contracts over 2 to 3 years. News Corp's OpenAI deal was reported at over $250 million across five years. Amazon reportedly pays the New York Times $20 million to $25 million a year, and OpenAI pays Dotdash Meredith at least $16 million a year. Most deals disclose no terms at all.
How much did Google pay Reddit for AI training data?
Reuters reported in February 2024, citing one source, that Google's deal with Reddit was worth about $60 million a year. Neither company confirmed it. Reddit's S-1 filing disclosed $203.0 million in aggregate data licensing contracts signed in January 2024, with 2 to 3 year terms and at least $66.4 million recognised in 2024, without naming the buyers.
How much is the News Corp OpenAI deal worth?
The Wall Street Journal, owned by News Corp, reported the May 2024 deal as potentially worth more than $250 million over five years, including credits for News Corp's use of OpenAI technology. That is roughly $50 million a year. Neither company disclosed financial terms. News Corp later signed a separate deal with Meta, reported at up to $50 million a year for about three years.
What does an AI training data licence cost per book?
The clearest public figure is the HarperCollins deal with Microsoft, reported at $5,000 per title, split evenly between the author and publisher, for selected nonfiction backlist titles. By comparison, the Anthropic copyright settlement pays roughly $3,000 per eligible work, which the court noted is four times the $750 statutory minimum for ordinary infringement.
Are AI licensing deals cheaper than copyright lawsuits?
In total, yes. Anthropic's $1.5 billion settlement exceeds any single reported licence. Per unit, not necessarily. The settlement works out near $3,000 per book, below the $5,000 HarperCollins licence fee. The difference is what each buys. A licence covers future use, display and provenance. The Anthropic settlement released only past acquisition and copying claims through August 2025.
Do AI licensing deals cover training or real-time retrieval?
Early deals focused on archives for training. Newer ones increasingly pay for retrieval, meaning current content shown or summarised inside AI products. Amazon's New York Times deal covers summaries in Alexa. Meta's 2025 news deals cover real-time news in Meta AI. Perplexity pays publishers by traffic, citation and agent use. Retrieval rights renew, which makes them more valuable to sellers.
Where to start if you own the content
Before any buyer calls, write down which grade of comparable you are anchoring on. If your target price comes from a Grade B leak, say so in your own model, and test what happens to your case if that number is half as large.
Then split your asset in two on paper: the archive you would sell once, and the current feed you would license by use. Price them separately. The archive sets your one-time windfall. The feed is the only part that can still be paying you in 2029.
Related analysis
If the asset you are pricing is your own research, read why original research is the content moat that holds. The legal detail behind the litigation rows is in the copyright cases that could reprice training data.
References
- Reddit, Inc., Form S-1 registration statement, 22 February 2024. Used for the $203.0M aggregate, 2 to 3 year terms, $66.4M 2024 minimum and buyer concentration. Figures located via SEC full-text search and read through TechCrunch's report of the filing. TechCrunch, Reddit says it's made $203M so far licensing its data, 22 February 2024. Used for the S-1 terms and the reported Google figure. Reddit, Inc., Form 10-K for fiscal 2025, 2026. Used for revenue, the licensing growth sentence and the two-partner concentration.
- Engadget, OpenAI will pay Dotdash Meredith at least $16 million per year, 19 November 2024, reporting Adweek's review of IAC financial documents. TheWrap, Amazon to pay New York Times $20 million a year, 30 July 2025, citing The Wall Street Journal. Media Copilot, Meta and News Corp AI content licensing deal, 4 March 2026, citing The Wall Street Journal of 3 March 2026.
- Authors Guild, Court grants final approval to Anthropic copyright settlement, 21 July 2026. Eligible works count of 482,460 from BetaNews coverage of the same order.
- CFO Dive, Wiley sees AI licensing revenue, 23 June 2025; and Publishers Lunch, Wiley expects AI sales to multiply, 18 June 2026.
- Bloomberg Law, Academic publishers sign AI deals as US cuts research grants, 10 June 2025. Used for Taylor & Francis and Informa figures. Sherwood News, HarperCollins will let Microsoft license some of its authors' books for AI, 21 November 2024, citing Bloomberg.
- Axios, Perplexity to give publishers 80% of new subscription revenue, 26 August 2025; and Axios, Disney to invest $1 billion in OpenAI, license characters for Sora, 11 December 2025.
- IPWatchdog, Third Circuit affirms fair use ruling against Ross, 30 September 2026; and DLA Piper, Getty Images v Stability AI, November 2025.
- The Decoder, Axel Springer and OpenAI licence worth tens of millions of euros per year, December 2023, reporting the FT; Bloomberg's three-year reading is reported by Verdict.
The weakest part of this source base: 8 of the 14 deal figures come from press reports citing unnamed sources, and two of those were reported by a newspaper owned by the seller. The Shutterstock figure rests on analyst reporting; its 10-K states only that data, distribution and services were 16% of 2023 revenue. Figures current as of 10 October 2026.
Related reading