From Shubhi K | Product & Market Analysis
Inference Costs Are Rising, Not Falling. The Software Margin Myth Is Breaking
On this page
Traditional software had near-zero marginal cost. Serving the ten-thousandth customer cost almost nothing more than serving the first. AI products do not work that way, and per-token prices falling steadily has not stopped total inference spend rising sharply. That gap between unit price and total cost is why AI-native gross margins sit below classic software, and why it matters to anyone buying it.
Key takeaways
- Falling unit prices have not produced falling bills. Per-token prices have declined steadily while total inference spend across the category has risen.
- Consumption grew faster than prices fell. Reasoning models and agent workflows consume many times the tokens a single chat exchange did.
- Cost of revenue in AI includes more than compute. Evaluation, retrieval infrastructure, monitoring and human review all sit inside the margin.
- The zero marginal cost assumption is what broke. Classic SaaS pricing was built on it, and AI pricing models are being rebuilt because it no longer holds.
The margin myth, and where it came from
Software became the most attractive business model in modern history for one structural reason. Once the product was built, distributing another copy cost almost nothing.
That produced gross margins in the seventies and eighties, which supported the entire venture and public market apparatus around SaaS. Growth was expensive, delivery was not, and every additional customer improved the ratio.
AI products break the second half of that. Every request consumes compute, and compute is metered. The tenth thousand user costs roughly what the first did, which is an entirely different business.
Why total cost rises while unit prices fall
Both statements are true simultaneously and the reconciliation is straightforward once stated.
A single question and answer in 2023 consumed a few thousand tokens. The same job today may involve retrieval across a document set, a reasoning trace, several tool calls and a verification pass. The unit got cheaper and the number of units grew by more.
The direction of travel at the model providers reflects this. Reported inference spending at one major lab rose from $3.76 billion across 2024 to $5.02 billion in the first half of 2025 alone, while headline per-token prices were falling throughout.
The capability cost curve makes it worse
Improving model quality is not linear in cost. A widely cited rule of thumb from market observers holds that making models twice as good costs roughly five times the energy and money.
If that holds even approximately, capability improvements arrive with cost increases attached, and those costs eventually reach the price list.
What agents changed
Agentic workflows are the single largest driver of consumption growth, and the mechanism is simple. An agent does not make one call. It makes a sequence, each conditioned on the last.
A task that a person would have done with one prompt becomes a plan, several tool invocations, error handling, a retry and a summary. Ten to fifty calls where there was one, each carrying the accumulated context of the ones before it.
Context accumulation is the compounding factor. Every step carries forward the history, so later calls in a chain are more expensive than earlier ones. Costs scale worse than linearly with task complexity.
This is why buyers who negotiated on per-token price find their bills unrecognisable a quarter later. The price they negotiated was correct. The volume assumption underneath it was not.
Retries and verification are pure cost
An agent that fails a step and retries has consumed the tokens for both attempts. A verification pass that confirms an output was correct has consumed tokens producing no new information.
Both are necessary for reliability and neither produces billable value in a per-outcome pricing model. That is why vendors selling outcome-based pricing carry more cost volatility than their pricing pages suggest, and why they cap usage more aggressively than their marketing implies.
What actually sits in AI cost of revenue
Compute gets all the attention and it is not the whole picture.
The classification question matters more than it sounds. Human review of AI outputs is a delivery cost in substance. Placing it in operating expenses rather than cost of revenue improves reported gross margin without changing anything real.
When comparing two AI companies' margins, check what each includes before concluding one is more efficient. Frequently the difference is accounting policy rather than operations.
What this means for pricing
Three consequences follow directly, and all three are already visible in the market.
Flat-rate pricing becomes dangerous. A heavy user on an unlimited plan can consume more than they pay, and the vendor's incentive shifts toward limiting the customers who get most value from the product.
Free tiers become expensive. The free tier worked because marginal cost was near zero. It is not, so free tiers are converting into trials and credit allocations across the category.
And usage-based pricing becomes close to inevitable, which shifts volatility onto the buyer. The pricing shifts this produces are examined further in the analysis of where AI returns are measured.
How this reaches your contract
The structural cost picture translates into four specific things buyers are now seeing at renewal.
Committed usage tiers replacing unlimited plans, with overage priced considerably above the committed rate. Rate limits that were previously generous being formalised into the contract rather than left as fair use. Credit systems that decouple headline price from actual consumption. And shorter contract terms, which give the vendor room to reprice as their own costs move.
None of these is unreasonable given the cost structure. All of them shift volatility from the vendor to you, and most buyers accept them without negotiating the one term that matters, which is a cap or a notice period on price changes.
If you take one thing into a renewal conversation, take that. A vendor facing genuine cost uncertainty will resist a hard cap and will frequently accept a notice period, which at least converts a surprise into a planning problem.
The case that this resolves itself
The counter-argument is serious and it may well be right.
Hardware efficiency is improving on a steep curve. Each accelerator generation delivers more useful work per unit of energy, and inference-specific silicon is arriving from several directions.
Model efficiency is improving alongside it. Smaller models increasingly match larger ones on defined tasks, and routing simple work to cheap models cuts cost substantially without a quality loss users can detect.
There is also a competitive floor. Providers competing for developer adoption have consistently priced below what a monopolist would charge, and that competition shows no sign of easing.
The honest position is that unit costs are falling and total costs are rising, and which one dominates depends entirely on whether consumption growth slows. Nobody has demonstrated that it is slowing.
What this means if you are buying, not investing
Most readers here are not deciding whether to own the equity. They are deciding what to standardise on and what to pay for it.
The practical implication of low switching cost runs in your favour. A category with weak lock-in is a category where annual contracts and aggressive pricing commitments are negotiable, because the vendor knows the alternative is a migration that costs your team a week rather than a quarter.
The second implication is that standardising has a real cost. Forcing a single tool on a team that runs several will meet resistance. The productivity argument for standardisation is also weaker here, because the tools genuinely do different jobs.
Most teams land on a middle path. They pay for one enterprise tool with the compliance and audit properties procurement requires, and tolerate individual use of others. That is more expensive than a clean standard and considerably cheaper than a rollout nobody adopts.
The security question is the one worth being firm on regardless. Any tool with repository access is a supply chain consideration, and personal accounts sit outside whatever controls you have built. Tolerating tool diversity is reasonable. Tolerating unmanaged repository access is not, and that distinction is where most policies should be drawn.
That framing also gives procurement something workable. Instead of a standardisation mandate nobody follows, use an approved list with defined access controls. Developers accept that, because it does not force them to abandon the tool they prefer.
What to watch
| Signal | Why it matters |
|---|---|
| Your own tokens per completed task | The only number that tells you whether efficiency gains are reaching you |
| Vendor gross margin disclosure and its definition | Margin figures are not comparable unless you know what each includes |
| Adoption of model routing | Routing simple work to cheaper models is the largest available cost lever |
| Changes to free tier or rate limits | The earliest visible sign a provider is under margin pressure |
The question to ask any AI vendor
One question separates vendors who have modelled their business from those who have not. What happens to our price if your inference costs rise by 40%?
A vendor with an answer has thought about the scenario and usually has a routing strategy, a caching layer or a contractual pass-through position. A vendor without one is carrying an unpriced risk that will eventually reach you, most likely at renewal and without warning.
The answer also tells you something about their gross margin. A vendor who says a 40% cost increase would be absorbed is telling you their margin is comfortable. A vendor who cannot answer is telling you it is not, whether or not they intend to.
Frequently asked questions
Are AI inference costs going up or down?
Both, depending on what you measure. Per-token prices have fallen steadily since 2023, while total inference spend has risen across the category. Consumption per task grew faster than unit prices fell, because reasoning models and agent workflows make many more calls, each carrying accumulated context. Your bill reflects total spend, not unit price.
Why did my AI bill increase when prices went down?
Because you are consuming far more units. A task that once took one call may now involve retrieval, a reasoning trace, several tool invocations and a verification pass. Each step carries forward the context of previous steps, so costs scale worse than linearly with task complexity. The negotiated price was correct; the volume assumption underneath it was not.
What are typical AI gross margins compared to SaaS?
Lower, though the comparison is complicated by inconsistent classification. Classic software achieved high margins because marginal delivery cost approached zero. AI products carry direct compute cost per request, plus retrieval infrastructure, evaluation, monitoring and often human review. Some companies classify review costs as operating expense rather than cost of revenue, which flatters the reported figure.
Will AI inference get cheaper over time?
Unit costs almost certainly will, through hardware efficiency, inference-specific silicon and smaller models matching larger ones on defined tasks. Whether total costs fall depends on whether consumption growth slows, and nothing currently suggests it is. The question is not whether efficiency improves but whether it improves faster than usage expands.
How can I reduce my AI costs?
Model routing is the largest available lever: send routine work to smaller, cheaper models and reserve frontier models for tasks where the quality difference is measurable. Then measure tokens per completed task rather than tokens overall, since that is the number that reveals whether efficiency work is reaching your bill or being absorbed by rising usage.
Why are free AI tiers disappearing?
Because free tiers were built on the assumption that marginal cost approached zero, which is not true for AI products. Every free request consumes metered compute. Providers across the category are converting free tiers into time-limited trials or fixed credit allocations, which is the most visible early sign of margin pressure in this business.
Where to start this week
One measurement, and it will probably surprise you.
Pick the AI workflow you run most often and calculate tokens consumed per completed task, not per request. Then compare it to the same figure three months ago if you have it. Most teams find consumption per task has grown considerably while they were congratulating themselves on falling unit prices.
That single ratio is the one to manage. It is the only number that tells you whether your costs are under control or simply being outrun by usage.
Then do one thing with the answer. Take it to your largest AI vendor and ask what they are doing about it on their side. The response tells you whether you are dealing with a company that has modelled its own economics or one that is hoping prices keep falling faster than usage grows.
References
- Kannan K R, The AI bubble: hype, capital and the trillion-dollar question, March 2026. Used for reported inference spending figures and the capability cost rule of thumb.
- Growth Unhinged, The 2026 state of B2B SaaS and AI monetization report, May 2026. Used for pricing model shifts across the category.
- The SaaS Library, B2B SaaS trends in 2026, May 2026. Used for the structural pricing analysis.
Inference spending figures are reported for individual companies and are not directly comparable across providers, since none disclose cost of revenue on a consistent basis. The relationships described here are structural rather than precisely measured.
Related reading