From Aryan Vatsa | Product & Market Analysis
AI Unit Economics and the Margin Trap: When Your Best Customer Loses You Money
On this page
Your heaviest user is often your least profitable one. AI product gross margins ran 45% in 2025, against the 70% to 80% that classic software taught the market to expect. The cause is not model pricing. It is AI unit economics under a flat rate. Cost tracks usage while revenue does not, so the customer getting the most value is the one draining the margin.
Key takeaways
- Flat-rate AI turns usage into an uncapped liability. Anthropic said one subscriber consumed tens of thousands of dollars of model usage on a $200 per month plan before weekly caps arrived in August 2025.
- The spread inside a customer base is wider than pricing assumes. Ramp's payments data put median monthly business AI spend at $2,246 against an average of $140,842, a gap of roughly 63 times.
- Contribution margin per account is the only number that finds the trap. A blended gross margin hides it, because profitable light users fund unprofitable heavy ones inside one reported percentage.
- The obvious fix is usually the wrong one. Heavy users are the reference accounts and the expansion pipeline, so capping them defends margin and damages the growth case at once.
This post is written for a founder or a director who owns the P&L on an AI feature. It picks that reader because the fix is a pricing decision, not a prompt engineering decision, and only one of those people signs it off.
The trap, stated as arithmetic
A flat-rate product charges a fixed price and absorbs variable cost. That works while the variable cost is small and bounded. Hosting a spreadsheet is both.
Model inference is neither. Every action carries a metered cost, and nothing in the product stops a determined user from taking a thousand actions a day. The price is capped. The cost is not.
Anthropic published the clearest example anyone has. When it introduced weekly rate limits in July 2025, it said some subscribers were running Claude Code continuously in the background, 24 hours a day. One user had consumed tens of thousands of dollars of model usage on a $200 plan. That is a contribution margin of roughly negative 10,000%, on one account, inside a product the market treats as a success.
Sam Altman said something similar about ChatGPT Pro in January 2025. He picked the $200 price himself, expected it to make money, and reported that OpenAI was losing money on it because people used it much more than expected. Two of the best-informed pricing teams in the industry both mispriced the same thing in the same direction.
Why blended gross margin hides it
Most teams look at one number: revenue minus cost of revenue, expressed as a percentage. That number is an average, and averages are exactly the wrong instrument here.
A book of business can report 68% gross margin while a fifth of its accounts lose money on every renewal. The light users subsidise the heavy ones, silently, until the mix shifts. Nothing in the reported percentage tells you when that shift is coming.
The instrument that works is contribution margin per account. It is not complicated, it is just not the number most dashboards were built to show. This is the same measurement gap that shows up on the cost side, covered in the piece on why inference costs sit inside AI gross margins differently to classic software costs.
What AI changed about the flat-rate bargain
Flat rate was never a gamble. It was a bet on a ceiling, and the ceiling was human attention.
A per-seat CRM could safely offer unlimited records because a salesperson can only type so much in a day. The heaviest user of a traditional tool might consume five times the median. That is a rounding error against a 75% margin.
The ratio in AI products is not five. Ramp's data on the businesses buying these products shows a 99th percentile spending about 370 times the median. Whatever the exact multiple is inside your own product, it is not five.
Agents removed the ceiling on effort
An agent works while the user sleeps. That single property breaks every assumption behind a flat rate, because consumption stops being a function of attention and becomes a function of ambition.
Cursor described the same effect from the vendor side. When it restructured its Pro plan in June 2025, it explained that new models spend more tokens per request on longer-horizon tasks, and the hardest requests cost an order of magnitude more than simple ones. Two users on identical plans can differ by 10 times on cost per request, before you count how many requests each one makes.
Multiply a 10 times cost-per-action spread by a 100 times actions-per-month spread and you have a customer base whose cost distribution spans four orders of magnitude, priced with one number. That is the trap. It is arithmetic, not mismanagement.
The contribution margin model you can copy
Here is the whole model. It fits in six columns and it will tell you more than any dashboard you currently own.
| Input | Symbol | Where you get it | Example |
|---|---|---|---|
| Monthly revenue per account | P | Billing system, net of discount | $50.00 |
| Non-AI variable cost | F | Hosting, support, payment fees, allocated per account | $10.00 |
| Billable AI actions per month | U | Product telemetry, per account, not averaged | varies |
| Blended cost per action | c | Provider invoice divided by total actions | $0.10 |
| Overhead multiplier | r | Retries, evaluation calls, failed runs, re-sent context | 0.25 |
| Contribution margin | CM | P minus F minus (U x c x (1 + r)) | computed |
The six inputs, and the one everybody gets wrong
Five of these are easy. The sixth, the overhead multiplier, is where most models understate the problem by a third.
Your provider bill is not your feature bill. It also contains retries after timeouts, evaluation and guardrail calls, context re-sent on every turn of a conversation, and runs the user abandoned halfway. None of that is visible in a per-action price list, and all of it lands in cost of revenue.
Set r honestly by dividing last month's total provider invoice by the count of actions your product logged as successful. If that ratio is 1.4, your r is 0.4. Do not guess it.
The breakeven usage line
Rearranged, the model gives you one number worth putting on a wall. Breakeven usage is (P minus F) divided by (c multiplied by (1 + r)).
With the illustrative values above, that is $40 divided by $0.125, which is 320 actions per month. Above 320 actions, the account destroys value. Below it, the account funds the company.
| Account type | Actions per month | AI cost | Contribution margin | Margin % |
|---|---|---|---|---|
| Median user | 40 | $5.00 | $35.00 | 70% |
| 90th percentile | 300 | $37.50 | $2.50 | 5% |
| Breakeven line | 320 | $40.00 | $0.00 | 0% |
| 99th percentile | 2,000 | $250.00 | minus $210.00 | minus 420% |
Every figure in this table is derived from the illustrative inputs in the previous table, not measured from a live product. The structure is the point. Substitute your own P, F, c and r and the shape holds, only the crossover moves.
The cross-subsidy ratio
Run the model across your base and one ratio falls out. How many median accounts does it take to fund one heavy account?
With the values above, it takes six. A heavy account at minus $210 needs six median accounts at $35 each to get back to zero. That sounds survivable, and it is, while heavy accounts are 1% of the base.
Now solve for the tipping point. With those inputs the whole book goes negative once heavy accounts pass 14.3% of customers. Adoption campaigns, agent launches and successful onboarding all push that share up. Your growth plan and your margin plan are pulling the same lever in opposite directions, which is the second-order effect nobody puts on a slide.
What the usage distribution actually looks like
Before modelling anything, get the shape. Most teams have never plotted it, and the plot is the argument.
Percentiles, not averages
Ramp's figures describe buyers rather than any single vendor's seats, so treat them as directional for your own product. The shape is what transfers, and the shape is the finding.
Pull the same five percentiles for your own accounts this week. If your median account uses 40 actions and your 99th uses 2,000, your pricing page is a bet you have never priced. That is also the mechanism behind the pattern described in falling token prices arriving alongside rising bills.
One warning about definitions. Count billable actions, not sessions or seats, and count them per paying account rather than per user. An enterprise account with 200 seats and three obsessive users is a heavy account, and per-user averages will never show you that.
What vendors do once they see the curve
Every serious AI vendor has now run this arithmetic, and the record of what they did is public. The pattern is remarkably consistent, and the dates tell you when each one hit the wall.
The specifics are worth having, because they show four different instruments. Anthropic used hard caps, adding weekly limits on 28 August 2025 that it said would touch under 5% of subscribers. Figma used metered seats, announcing in December 2025 that full seats carry 3,000 to 4,250 credits a month and that limits would be enforced from 18 March 2026. Microsoft used a second currency, moving agent billing to Copilot Credits from 1 September 2025 at a pay-as-you-go rate of $0.01 per credit. GitHub used complexity multipliers, pricing extra premium requests at $0.04 each and assigning code review a multiplier of 13.
| Response | What it protects | What it costs | When I would use it |
|---|---|---|---|
| Hard caps and weekly limits | Margin, immediately and completely | Trust with your loudest advocates, who are the capped ones | Only against outlying abuse, resale, and unattended background automation |
| Credits metered into the seat | Predictability on both sides of the contract | A second currency the buyer has to learn and forecast | Products with one countable unit of work a buyer recognises |
| Flat base plus usage overage | Alignment between what you charge and what you spend | Budget anxiety, which slows adoption of the feature you want adopted | Sophisticated buyers who already run cost controls on cloud |
| Silent throttling or model downgrade | Margin, with no pricing announcement | The relationship, on the day a customer benchmarks you and finds out | Never as a policy. Only as a disclosed degradation mode with a status page |
| Where caps genuinely win | Nothing else stops a single account from consuming unbounded compute | Nothing, if the cap is set above the 99th percentile of legitimate use | Any product where an agent can run unattended overnight |
My position is that the fourth row is the only unacceptable one. Caps, credits and overage are all defensible pricing choices that a buyer can plan around. Quietly routing a paying customer to a cheaper model is a product change sold as a service level, and it is the fastest way to turn a pricing problem into a credibility problem.
The market has already moved in the same direction. ICONIQ found that among the AI builders it surveyed, consumption-based pricing rose from 35% to 42% adoption and outcome-based pricing from 18% to 23% in six months, with companies blending 1.7 models on average. That is the same conclusion reached in the analysis of why per-seat pricing stopped matching delivered value, arriving from the cost side rather than the value side.
Where this argument is weakest
Three genuine objections, and the third one is the strongest.
Costs per unit of work are falling
The same ICONIQ survey reports gross margins rising from 45% in 2025, with respondents expecting further improvement as they route most workloads to smaller or fine-tuned models and escalate only hard tasks. Two-thirds of the companies surveyed said per-query unit economics had improved.
If that continues, the breakeven line moves right every quarter and today's unprofitable account becomes next year's profitable one. A vendor that panics and caps hard in 2026 may have priced against a cost curve that no longer exists in 2027. That is a real risk and I would weight it seriously before writing caps into a contract.
The customer you would never fire
The heavy user is not a parasite. They are usually the account with the highest retention, the deepest workflow integration, the best case study and the largest expansion potential. Contribution margin at the account level is a one-period measure, and it says nothing about lifetime value.
Firing them, or capping them into leaving, can be the most expensive margin improvement you ever book. The honest version of this analysis compares contribution margin against acquisition cost and retention, not against zero. I would tolerate a negative-margin account that renews for four years and closes two referrals, and I would say so out loud in the pricing review.
There is also a survey limitation worth stating plainly. The ICONIQ margin figures are self-reported by roughly 300 mostly sub-$100 million companies, and self-reported margin definitions vary in what they put above the line. Treat 45% as directional evidence of a category, not as an audited benchmark.
If you are on the buying side of this
The same arithmetic runs in reverse, and it tells you what your renewal will look like.
Any vendor selling you unlimited AI on a flat seat is either capping it later or losing money on you now. Neither is stable. The useful question in a sales conversation is not what the price is, it is what happens when your team doubles its usage.
Ask for three things in writing. The included allowance expressed in a unit you can measure, the overage rate once it is exhausted, and the notice period before allowances change. A vendor that will commit to all three has done this model. A vendor that answers with the word unlimited has not, and you are looking at a repricing in the next 18 months.
Then instrument your own side. Most buyers cannot say which of their teams generated 80% of last month's AI actions, which makes negotiating an allowance guesswork. The discipline is the same one described in the analysis of where measurable AI return has actually appeared. It applies just as well to the build against buy decision on coding agents, where the vendor's cap becomes your ceiling.
Frequently asked questions
Why do AI features reduce SaaS gross margin?
Because every AI action carries a metered compute cost, while classic software had near-zero marginal cost per user. Serving one more request costs real money in a way that serving one more spreadsheet did not. AI product companies surveyed by ICONIQ reported average gross margins of 45% for 2025, against the 70% to 80% range that traditional software businesses have historically targeted.
How do you calculate contribution margin per customer for an AI product?
Take monthly revenue for the account and subtract non-AI variable costs such as hosting and support. Then subtract billable AI actions multiplied by your blended cost per action and an overhead multiplier for retries, evaluation calls and re-sent context. The result is contribution margin. Divide by revenue for a percentage. Run it per account, never as a blended average across the base.
What is the power user problem in AI pricing?
It is the gap between what heavy users pay and what they cost. Under a flat rate, cost scales with usage but revenue does not, so the most engaged customer can carry a negative margin. Anthropic cited one subscriber consuming tens of thousands of dollars of model usage on a $200 per month plan before it introduced weekly rate limits in August 2025.
Is flat rate pricing dead for AI software?
Not dead, but rarely offered without a limit behind it. Vendors keep a flat headline price and attach an allowance, a credit balance or an overage rate. Figma enforced seat credit limits from March 2026, GitHub charges $0.04 per additional premium request, and Microsoft bills agents in Copilot Credits. The list price stays, the meaning of unlimited does not.
How many AI actions can a $50 seat cover?
It depends entirely on your cost per action, which is why you calculate it rather than assume it. With $10 of non-AI variable cost and an all-in cost of $0.125 per action, a $50 seat breaks even at 320 actions a month. Halve the cost per action and the line moves to 640. That single number should govern your allowance design.
What should buyers do about AI usage caps?
Get the allowance, the overage rate and the notice period in writing before signing. Measure your own usage by team so you can forecast against the allowance rather than guessing. Assume any vendor offering unlimited AI on a flat seat will reprice within roughly 18 months, and negotiate the terms of that repricing now, while you still have leverage in the deal.
Where to start this week
One measurement, one decision, in that order.
The measurement takes an afternoon. Export last month's provider invoice, divide it by the count of successful actions your product logged, and you have c and r together. Then pull actions per paying account and sort descending. The top 1% of that list is your exposure, and you can read the answer off the screen.
The decision is what to do about the accounts above your breakeven line. I would not cap them first. I would call the top ten and ask what they are using it for. The question is whether you are looking at a pricing failure or an underpriced product tier that those customers would happily pay for. In my experience the second is more common than founders expect, and it is the only response to this trap that increases revenue instead of defending it.
Related analysis
This piece works from the account level up. For the category-level view, read what actually sits inside AI cost of revenue, and for the value-side version of the same break, read how vertical AI is repricing horizontal software.
References
- ICONIQ Growth, State of AI 2026: The Builder's Economy, July 2026. Survey of about 300 executives in Q2 2026. Used for gross margin figures and pricing model adoption.
- Ramp, How much do AI tokens cost businesses, 2026 spending benchmarks, 8 June 2026. Used for all spend percentiles and the median against average gap.
- TechCrunch, Anthropic unveils new rate limits to curb Claude Code power users, 28 July 2025. Used for the weekly caps, the under 5% figure and the single-user cost example.
- Cursor, Clarifying our pricing, 4 July 2025. Used for the order-of-magnitude spread in cost per request and the $20 included usage change.
- Figma, Updates to AI credits in Figma, 9 December 2025. Used for seat credit allowances and the 18 March 2026 enforcement date.
- GitHub, Requests in GitHub Copilot, accessed August 2026. Used for the $0.04 premium request rate and the code review multiplier.
- Microsoft Learn, Purchase and manage Copilot credits, accessed August 2026. Used for the $0.01 pay-as-you-go credit rate.
- Fortune, Sam Altman says OpenAI is losing money on ChatGPT Pro, 7 January 2025. Used for the Altman comment on the $200 plan.
The weakest part of this source base is the model itself. Every figure in the worked example is illustrative, chosen to be arithmetically clean rather than measured from a live product, and the ICONIQ margins are self-reported by a survey panel rather than audited. The vendor pricing dates and the Ramp percentiles are the load-bearing facts here. The spreadsheet is a structure to fill with your own numbers.
Related reading