From Aryan Vatsa | Product & Market Analysis

AI Shrinkflation: Your Price Held, Your Allowance Did Not, and Nobody Told You

On this page

Your AI subscription has a price you can read and an allowance you cannot. That asymmetry is the whole story. When Anthropic added weekly rate limits to Claude Pro and Max in 2025, it gave subscribers 31 days of notice. Its published commitment to developers is at least 60 days before an API model is retired. Same company, same compute, two very different duties of disclosure.

Key takeaways

  • The unit of sale is the problem, not the size of the allowance. Cursor replaced a fixed 500-request monthly cap with $20 of usage billed at model API rates, which covered about 225 Sonnet 4 requests or about 650 GPT 4.1 requests depending on which model you chose.
  • Developers get a written notice policy. Subscribers get a blog post. OpenAI documents at least 6 months of notice before retiring a generally available API model. Neither OpenAI nor Anthropic publishes an equivalent floor for consumer plan allowances.
  • Price stability is doing the marketing work. GitHub moved every Copilot plan from premium requests to token-metered AI Credits on 1 June 2026 while holding Pro at $10 and Pro+ at $39.
  • Limits also went up in 2026, which is the strongest evidence against the shrinkflation reading. Anthropic doubled Claude Code 5-hour limits in May 2026 and ran a 50% weekly uplift into July. Competition moves allowances in both directions.
Under 5%Share of Claude Pro and Max subscribers Anthropic estimated its new weekly limits would affect. Source: Anthropic, reported by TechCrunch, 28 July 2025.
225Sonnet 4 requests covered by the $20 Cursor Pro plan after June 2025, replacing a fixed 500-request cap. Source: Cursor, 4 July 2025.
60 daysMinimum notice Anthropic commits to before retiring a public API model. No equivalent floor is published for subscription limits. Source: Anthropic docs, 2026.

What AI shrinkflation actually is

AI shrinkflation is a reduction in what a subscription delivers while the headline price stays the same. It arrives as a smaller allowance, a change to the unit the allowance is counted in, or the removal of a model you had built a workflow around. The price tag is stable, so the increase never shows up as an increase.

The grocery analogy is useful and it is also where most coverage stops. A chocolate bar has a declared weight printed on the pack. Shrink the bar and the printed number changes, which is why a regulator can write a disclosure rule about it.

Software subscriptions sold in messages, credits, sessions or requests have no declared unit in that sense. There is a limit, it is real, and in most cases the vendor has never published what it equals in tokens. You cannot compare this month's bar to last month's when neither one has a weight on it.

The interesting claim is not that vendors are charging more. It is that the pricing model in use makes the question unanswerable by the buyer, and unanswerable questions do not get raised at renewal.

The three mechanics, ranked by how hard they are to see

Almost every case filed under shrinkflation is one of three distinct things. Separating them matters, because only one is a genuine reduction and the other two are transfers of risk.

1. The allowance cut

The most visible version. A number goes down, or a new ceiling appears where none existed. Anthropic's weekly rate limits are the clearest documented example. Two new limits reset every 7 days, layered on top of the existing 5-hour caps, effective 28 August 2025 across the $20 Pro plan and both Max tiers.

A cap you can name is a cap you can plan around, and Anthropic announced it publicly rather than shipping it silently.

2. The unit swap

Harder to see and, in my view, the more consequential of the three. The allowance stops being counted in something you can count and starts being counted in dollars of metered consumption.

Cursor is the textbook case. Before 16 June 2025 the Pro plan carried a limit of 500 requests per month, with Sonnet models costing two requests each. After the change the same $20 bought unlimited use of Tab and Auto, plus $20 of frontier model usage at API pricing. Cursor's own worked example put that at about 225 Sonnet 4 requests, 550 Gemini requests, or 650 GPT 4.1 requests.

Under the old scheme your monthly ceiling was a constant. Under the new one it moves with model choice, context length and whatever the underlying API charges next quarter. The vendor stopped absorbing that variance and you started absorbing it.

3. The model swap

The version that produces the most noise and the least measurable loss. A model you relied on disappears from the picker, or requests quietly route to a cheaper sibling once you cross a threshold.

OpenAI removed GPT-4o from ChatGPT at the GPT-5 launch in August 2025, restored it after user objection, then retired it along with several older models on 13 February 2026. The company said the vast majority of usage had already shifted, with roughly 0.1% of users still selecting GPT-4o daily. A Change.org petition still gathered more than 20,000 signatures.

The removal was defensible on usage data and it still broke real workflows for people whose prompts had been tuned against that specific model.

How much warning you get depends on what you bought Days between public announcement and the change taking effect. Bars to scale. OpenAI API, GA model 180 days Anthropic API, Sonnet 4 62 days Copilot billing switch 35 days Claude weekly limits 31 days Cursor Pro plan change 0 days explained 18 days afterwards. Blue: written policy floor. Red: one-off announcement with no published floor behind it.
The top two bars come from published deprecation policies. The bottom three are individual decisions, and nothing obliged any of them to be longer.

What actually changed, vendor by vendor

Four documented changes, all with the headline subscription price held constant.

Documented allowance changes at an unchanged price, 2025 to 2026
Vendor and planBeforeAfterPrice
Cursor Pro, June 2025500 requests per month, Sonnet counted as 2$20 of usage at API rates, about 225 Sonnet 4 requests$20, unchanged
Claude Pro and Max, August 20255-hour rolling session limitsTwo additional weekly limits, one overall and one for the top model$20, $100, $200, unchanged
GitHub Copilot Pro, June 2026300 premium requests, $0.04 each beyond$10 of AI Credits metered on input, output and cached tokens$10, unchanged
ChatGPT, February 2026GPT-4o selectable in the pickerGPT-4o retired with several older modelsPlan prices unchanged

The Copilot Pro+ tier followed the same pattern at $39 per month with $39 of credits. Users on existing annual plans stayed on legacy request-based billing until expiry.

Cursor: the request that stopped being a request

The direction of travel matters more than the arithmetic. On the old plan, 500 request units with Sonnet at two each gave you roughly 250 Sonnet turns, and 250 was 250 in January and in December.

On the new plan the same $20 buys about 225 Sonnet 4 requests. That is a modest reduction on that specific model. The larger change is that the number is now derived rather than declared, so it moves when model prices move and when your context windows grow.

Cursor's chief executive apologised on 4 July 2025 for unclear communication and offered refunds for unexpected charges between 16 June and 4 July. It confirms the users were not imagining the change.

GitHub: a credit is not a request

GitHub announced on 27 April 2026 that premium request units would be replaced by AI Credits from 1 June 2026. The new plans hold Pro at $10 per month including $10 of credits and Pro+ at $39 including $39. Usage is now calculated on tokens consumed, including input, output and cached tokens, at each model's listed API rate.

Run the old numbers next to the new ones. Copilot Pro included 300 premium requests with additional requests at $0.04. Priced at GitHub's own overage rate, that allowance was worth $12. The replacement budget is $10.

I would not lean on that comparison too hard. An overage price is not the same thing as the value of an included unit, and a short cheap request under the new scheme can cost far less than $0.04. What the comparison does establish is that the two schemes are not obviously equivalent, and GitHub has not published a conversion that would let a buyer check.

Same $20. The ceiling now depends on which model you pick. Monthly requests included in the Cursor Pro plan, before and after 16 June 2025. BEFORE, ONE FIXED NUMBER AFTER, THREE DIFFERENT NUMBERS 250 Sonnet turns from 500 units 225 Sonnet 4 550 Gemini 650 GPT 4.1 Source: Cursor, 4 July 2025. The 250 figure applies Cursor's stated rule that Sonnet cost two of the 500 units.
The red bar is the only like-for-like comparison here. The two blue bars are the real change: your allowance is now a function of a choice you make inside the product.

The notice gap is the finding

Here is the part I did not expect when I started checking. The same companies that give subscribers a few weeks of warning publish precise, binding notice policies for developers, and they honour them to the day.

OpenAI's API deprecation policy commits to at least 6 months for generally available models. It gives at least 3 months for specialised variants such as Codex and deep research models, and as little as 2 weeks for anything carrying preview in the name. Anthropic's documentation states that it provides at least 60 days notice before model retirement for publicly released models, with a four-stage lifecycle of active, legacy, deprecated and retired.

What 60 days actually bought

Check the arithmetic rather than the promise. Anthropic deprecated Claude Sonnet 4 and Claude Opus 4 on 14 April 2026 and retired both on 15 June 2026, which is 62 days. Claude Haiku 3 was deprecated on 19 February 2026 and retired on 20 April 2026, exactly 60. The floor is real and it is being run close to.

Now compare the subscription side. Anthropic's weekly limits were announced on 28 July 2025 and took effect on 28 August 2025, a gap of 31 days. GitHub announced the Copilot billing change on 27 April 2026 for 1 June 2026, a gap of 35 days. Cursor changed the Pro plan on 16 June 2025 and published the explanation on 4 July.

Notice periods, developer contract versus consumer plan
ChangeAnnouncedEffectiveNotice
OpenAI API, generally available model retirementPer published policyVariesAt least 180 days
Anthropic API, Claude Sonnet 4 and Opus 4 retirement14 April 202615 June 202662 days
GitHub Copilot, premium requests to AI Credits27 April 20261 June 202635 days
Claude Pro and Max, new weekly rate limits28 July 202528 August 202531 days
Cursor Pro, request cap to usage budgetExplained 4 July 202516 June 2025None in advance

The first two rows are backed by written policies that a buyer can cite. The last three are individual decisions. That is the distinction, not a judgement about whether 31 days was enough.

Developers get a contract because their integration code breaks loudly and publicly. Subscribers get a changelog because a subscriber who runs out of allowance simply waits, and waiting is invisible to everyone except the person doing it.

Why vendors are doing this, in one paragraph of economics

A flat monthly fee against a variable marginal cost only works while heavy users stay rare. Inference is not free at the margin, and agentic tools changed the shape of the usage curve by turning one human instruction into hundreds of model calls. Anthropic said as much when it introduced weekly limits, citing users running Claude Code continuously in the background and describing itself as very constrained on compute.

The $20 price point is the second half of the explanation. It became the category anchor early, and moving off it is a public act that invites comparison. Adjusting an allowance is not. Given a choice between changing a number everyone can see and changing a number almost nobody measures, any rational pricing team changes the second one.

None of that requires bad faith. I think it is a pricing model the vendors cannot yet commit to, being sold in a format that implies they can. The unit economics behind that gap are covered in the piece on how inference costs run through AI gross margins. The same pressure is reshaping headcount-based pricing in the analysis of seat compression across SaaS.

Where this argument is weakest

Three genuine problems with everything above. The first is the one that nearly killed the post.

Limits went up in 2026, not down

Anthropic permanently doubled the Claude Code 5-hour limits for Pro, Max, Team and seat-based Enterprise plans in May 2026. It then ran a promotion giving 50% higher weekly limits, extended to 19 July 2026. Prices did not change. That is inflation running backwards, at the same vendor that introduced the weekly caps.

This is the strongest evidence against my own argument and I am not going to bury it. What it shows is that allowances are a competitive instrument, not a ratchet. When a rival ships a credible coding agent, the allowance goes up. The asymmetry of disclosure survives either way. The uplift was announced with the same advance notice as the cut: not much, and with an end date attached.

Constraint is a real explanation, not an excuse

Compute has been supply-limited for three years. A vendor that lets a small number of accounts consume a disproportionate share is degrading service for everyone else, and rationing is the ordinary answer to that. Anthropic's stated target was account sharing, reselling and continuous background use, which is not a description of your team.

If you want the scale of the constraint rather than the rhetoric, the capex piece lays out what the industry is spending to relieve it.

Nobody can produce the counterfactual

The honest hole in every shrinkflation claim, including this one, is that no independent party measures what a plan delivered last quarter against what it delivers now. Anthropic's under 5% estimate cannot be checked because the underlying usage distribution is not published. I would treat that figure as unfalsifiable rather than false, which is a different and more useful complaint.

The same applies to the widely repeated claim that models get quietly quantised or downgraded under load. I found no primary evidence for it, and I am not going to assert it on the strength of restatement volume. Six blog posts citing each other is one source.

Which changes you can actually catch Dark blue: yes. Pale: partial or vendor dependent. Red: no. Announced ahead Shows on invoice Measurable by you Allowance cut Usually No Yes Unit swap Yes No Only with a baseline. Model swap API yes, app varies. No Yes Silent routing Rarely No Weak evidence only. The invoice column is red four times out of four. That is the design feature, and it is why the receipt never argues with you.
Read the middle column. A flat price means the invoice is constant by construction, so it can never be the thing that alerts you.

How to detect it in your own account

You cannot audit a vendor's allowance. You can audit your own consumption, and that is enough. A change in the plan shows up as a change in what a constant task costs you.

Run a canary task. Pick one job your team does every week that is genuinely repeatable. The same repository, the same document, the same prompt, the same model, run on the same day. Log three things each time: how long it took, how much of the allowance it consumed, and whether it completed on the first attempt. Ten weeks of that data settles arguments that no amount of forum discussion will.

Record the plan terms as text, not as a memory. On the day you buy, copy the vendor's current limits page into your own notes with the date. Vendors update documentation in place. If you do not keep a dated copy, you will be arguing from recollection against a page that has silently changed.

Watch the fallback, not the failure. Modern tools rarely refuse. They downgrade. The visible signals are a switch to a mini or lightweight model after a threshold, a longer queue at peak hours, a shorter effective context, or a sudden willingness to summarise instead of read. Instruct your team to report the downgrade, because most people quietly work around it and never mention it.

Track cost per completed unit of work, not cost per seat. A seat price that holds while the number of tickets closed per seat falls is a price rise. This is the same discipline argued in the piece on where measurable AI return has actually shown up. It is the only version of the measurement you can perform without vendor cooperation.

What to write into the renewal

Everything above is diagnosis. This part changes an outcome, and it only works before you sign.

Ask for the allowance to be expressed in a unit you can count, and get the conversion in writing. If the plan is sold in credits, the vendor knows the token rate behind them, because they compute it to bill you. A supplier who will not state the conversion is telling you the conversion is expected to move.

Then ask for two clauses. The first is a notice floor on material changes to included usage, pegged to something the vendor already publishes. Both OpenAI and Anthropic have written API deprecation policies, so the ask is not novel and it is awkward to refuse. The second is a model continuity clause: advance notice when the model underneath the product is retired, replaced or re-routed. That matters because your prompts are tuned to a specific model whether or not anyone wrote that down.

Neither clause is standard on a $20 self-serve plan and you will not get them there. Both are entirely gettable on an annual team or enterprise agreement, and almost nobody asks. The same negotiating window is where the wider rationalisation work pays off, and if you are still choosing between vendors, the build versus buy comparison for coding agents covers what else to test before committing.

One more thing, which is the counterintuitive part. I would rather have a smaller allowance stated in a unit I can count than a larger one stated in a unit the vendor defines. Certainty is worth paying for, and it is currently underpriced by buyers who optimise the headline number.

Frequently asked questions

What is AI shrinkflation?

AI shrinkflation is a reduction in what a subscription delivers while the headline price stays the same. It shows up in three ways: a smaller usage allowance, a change in the unit the allowance is counted in, or the removal of a model you were relying on. Groceries have a declared weight on the pack. Most AI plans do not declare a comparable unit, so the reduction is hard to prove.

Are AI companies reducing usage limits without telling users?

Mostly they announce the change, then define it in units you cannot audit. Anthropic publicly announced weekly limits for Claude Pro and Max on 28 July 2025, 31 days before they took effect, and estimated they would affect under 5% of subscribers. GitHub announced its move from premium requests to AI Credits 35 days ahead. The announcement is real. The measurement you would need to check it is not published.

How do I tell if my AI subscription limits have been cut?

Run a fixed task on a fixed schedule and record what it consumes. Use the same prompt, the same file and the same model every week, then log how much of your allowance it burns and how long it takes. A vendor can change an allowance quietly. It cannot hide a repeatable task that starts costing 30% more of your quota to complete.

Why do AI subscriptions have usage limits at all?

Because inference has a marginal cost per request and a flat monthly fee does not. Anthropic said in July 2025 that it was very constrained on compute while adding weekly limits aimed at users running Claude Code continuously. A flat price with unlimited use is a bet that heavy users stay rare. Limits are how a vendor stops that bet from going wrong.

Do AI vendors have to give notice before changing usage limits?

No general rule requires it, and the gap between what vendors promise developers and what they promise subscribers is wide. OpenAI documents at least 6 months of notice before retiring a generally available API model. Anthropic commits to at least 60 days for publicly released models. Neither publishes an equivalent commitment for subscription usage allowances, which are governed by terms the vendor can revise.

Is AI shrinkflation illegal?

Not on current evidence. France has required supermarkets over 400 square metres to flag reduced pack sizes since 1 July 2024, and the EU has no dedicated shrinkflation law, leaving it to general unfair practices rules. Those regimes assume a declared unit of measurement. Software subscriptions sold in messages, credits or sessions do not have one, so the disclosure duty has nothing to attach to.

Where to start this week

Open the billing page of your largest AI subscription and find the sentence that states how much usage is included. If that sentence contains a unit you cannot convert into tokens or requests, you have found the exposure, and you have found it before renewal rather than during an outage.

Then set up the canary. One repeatable job, one fixed prompt, one fixed model, logged weekly in a spreadsheet nobody has to maintain. It costs an hour to build and it is the only instrument that will ever tell you whether the thing you are paying for got smaller. Vendors will keep adjusting allowances, in both directions, because compute economics force them to. The buyers who notice will be the ones who wrote down what normal looked like.

Related on pricing mechanics

Allowances are one lever. The others are covered in seat compression across SaaS pricing and in what inference actually costs the vendors setting these limits.

References

  1. TechCrunch, Anthropic unveils new rate limits to curb Claude Code power users, 28 July 2025. Used for the weekly limit announcement, the 28 August 2025 effective date, the under 5% estimate and the compute constraint comment.
  2. Cursor, Clarifying our pricing, 4 July 2025. Used for the 500-request cap, the Sonnet two-request rule, the 225, 550 and 650 request figures, the apology and the refund window.
  3. GitHub Blog, GitHub Copilot is moving to usage-based billing, 27 April 2026. Used for the 1 June 2026 effective date, the AI Credit allowances and the unchanged $10 and $39 prices.
  4. GitHub Docs, Requests in GitHub Copilot (legacy), retrieved August 2026. Used for the 300 premium request allowance and the $0.04 per additional request rate.
  5. Anthropic, Model deprecations, retrieved August 2026. Used for the 60-day notice commitment, the lifecycle labels and the Sonnet 4, Opus 4 and Haiku 3 deprecation and retirement dates.
  6. OpenAI, Deprecations, retrieved August 2026. Used for the 6-month, 3-month and 2-week API notice periods.
  7. Help Net Security, Claude Code users keep 50% higher limits until July 19, 13 July 2026. Used for the 2026 limit increases and the promotion end date.
  8. Bird & Bird, Shrinkflation in France: new obligation to inform consumers, 2024. Used for the 1 July 2024 start date and the 400 square metre threshold.

The weakest thing about this source base: every allowance figure comes from the vendor that set it, and no independent party measures delivered usage against advertised limits over time. The GPT-4o retirement details and the February 2026 usage share are drawn from press coverage of OpenAI's announcement rather than from the announcement page, which did not resolve when retrieved on 20 August 2026. Figures are current as of that date and allowances in this category change within weeks.

SK
Aryan Vatsa
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading