From Shubhi K | Product & Market Analysis
AI Billing Units Are Now a Sales Blocker, Not a Finance Complaint
On this page
Every AI vendor invented its own billing unit, and none of them convert. A Copilot Credit, a Flex Credit, an outcome and an automated resolution are four different objects carrying four different prices. AI billing has quietly become the slowest part of an AI purchase. 16 to 20 weeks is now the most common buying cycle, against 7 to 10 weeks for ordinary software.
Key takeaways
- The unit is the problem, not the price. Microsoft bills Copilot Credits at $0.008 each, Salesforce bills Flex Credit actions at $0.10, Intercom bills outcomes at $0.99 and Zendesk bills automated resolutions at $1.50. None of those four units describes the same event.
- Inside a single vendor, the spread is 100 to 1. Copilot Studio charges 1 credit for a classic answer and 100 credits for a premium generative tool response. Which one your agent fires is a design decision the buyer never sees.
- Finance is now in the room and it is saying no. Finance involvement in software buying rose from 31% to 46% in a year, and 49% of buyers had an already-approved purchase vetoed by the CFO.
- Having a token budget correlates with more vetoes, not fewer. Among organisations with a dedicated token or LLM budget, 54% reported a CFO block, against 29% of those without one. Visibility surfaces the variance rather than removing it.
What a credit actually is, and why no two vendors agree
A controller can audit a seat. It maps to a person, a payroll record and a leaver process. A credit maps to nothing outside the vendor that issued it.
That is the whole story, and everything below is the consequence.
A credit is a private currency
Underneath almost every AI product sits a token meter run by a model provider. Tokens are at least comparable. A thousand input tokens is the same object at every lab, priced differently.
Vendors do not resell tokens. They wrap tokens, orchestration, retrieval, retries and margin into a unit they name themselves, then price that unit. The wrapping is legitimate. The naming is where the trouble starts, because each vendor picks a different noun and a different denominator.
Scott Grossman, CFO of Ensono, told diginomica in August 2026 that there is "no perfect way to forecast AI spend, because frankly, we're just learning what goes into the cost of a token". His team has had to build its own translation layer as vendors shift reporting from tokens to credits mid-contract.
Nothing converts between them
You cannot express a Copilot Credit in Flex Credits. You cannot express either in resolutions. There is no exchange rate, no clearing house, and no vendor with an incentive to publish one.
That absence does real work in a negotiation. When every account team quotes in its own currency, every account team gets to frame its currency as cheap. A buyer comparing $0.008 against $2.00 is not comparing prices at all. They are comparing a rounding error against a whole customer interaction.
Five vendors, five units, one consolidated bill
Here is the same category, priced five ways. Every rate below comes from the vendor's own documentation, read in August 2026.
| Vendor and unit | Published rate | What triggers the charge |
|---|---|---|
| Microsoft Copilot Credit | $0.008 in a 25,000 credit pack at $200 a month. | Any billable agent event, from 1 to 100 credits depending on type. |
| Salesforce Flex Credit action | $0.10 per standard action. | Each function the agent executes on the platform. |
| Salesforce Agentforce conversation | $2.00 per conversation. | The session, whether or not it succeeds. |
| Intercom Fin outcome | $0.99 per outcome. | Customer confirms resolution, stops asking, or a workflow completes. |
| Zendesk automated resolution | $1.50 per automated resolution. | AI handles it with no human, and the ticket stays closed for 72 hours. |
Rates are list prices from vendor pricing pages and documentation. They are not comparable to each other, which is the point of the table rather than a flaw in it. Contracted rates, volume tiers and overage multipliers are negotiated and mostly unpublished.
The same question can cost 1 credit or 100
Cross-vendor confusion is the visible half. The harder half sits inside one vendor's own rate card.
The 1-to-100 spread inside one product
Microsoft publishes its Copilot Studio rates in full, which makes it the fairest available example rather than the worst offender. A classic answer costs 1 credit. A generative answer costs 2. An agent action costs 5. Tenant graph grounding costs 10. A premium text and generative AI tool costs 100 credits per 10 responses.
A buyer sizing a deployment is asked to forecast a number that depends on which of those paths an agent takes. That path is set by how the agent was configured, which knowledge sources it reaches for, and whether the orchestrator decided a step was needed. None of it is visible from the outside at signature.
Microsoft's own worked example makes the compounding clear. A support agent averaging four classic answers and two generative answers across 900 daily conversations runs at 7,200 credits a day. Add tenant graph grounding to the same agent and the per-response cost moves from 2 credits to 12.
Reasoning models bill on two meters at once
When an agent calls a reasoning-capable model, Copilot Studio bills the feature rate plus a premium tool rate of 10 credits per 1,000 tokens. Two meters, one response.
This is the detail that breaks a spreadsheet. A model upgrade that a platform team treats as a quality improvement arrives in finance as a second billing line nobody forecast. The analysis of what inference costs do to AI margins covers the supply side of the same pressure.
Outcome pricing moved the argument, it did not settle it
The obvious fix is to stop billing effort and start billing results. Several vendors have done exactly that, and it is a genuine improvement.
It also created a second definitional problem on top of the first.
Three definitions of one word
Zendesk bills an automated resolution when an AI agent handles an interaction with no human involvement and the customer does not reopen the ticket within a 72 hour quiet period. Intercom bills an outcome when the customer confirms resolution, when they stop asking after a reply, or when a workflow completes. Intercom states that handoffs count.
Those two definitions differ on the case that matters most commercially. Consider an interaction the agent partly handles before passing it to a human. One vendor bills it. The other does not. Both call the result a resolution.
My position, stated plainly, is that outcome pricing is better than session pricing and should not be sold as certainty. It converts a quality risk into a volume risk. Finance can budget for volume more easily than for model behaviour, so this is progress. It is not the predictability the marketing implies.
| Meter | Buyer carries | Vendor carries |
|---|---|---|
| Per seat | Underuse. You pay for licences nobody opens. | Nothing on volume. The bill is fixed. |
| Per credit or token | Everything. Model choice, retries and agent design all land on the buyer. | Nothing. Cost of goods is passed through with margin. |
| Per session or conversation | Failure. An unresolved session bills the same as a resolved one. | Depth. A long session costs the vendor more at the same price. |
| Per outcome or resolution | Volume, and the definition of success. | Quality. Failures are unbilled work. |
What this actually does to a forecast
Seat maths is arithmetic. Headcount times price, adjusted for a hiring plan you already own.
Consumption maths is a distribution. You need a volume forecast, a units-per-interaction estimate, an overage rate and a view on how the vendor's own product changes before renewal. Three of those four are outside the buyer's control.
The FinOps discipline has felt this arrive at speed. 98% of practitioners now manage some form of AI spend, against 31% two years earlier, and the survey names visibility, allocation and value measurement as the three main obstacles. The most requested capability is granular monitoring of tokens, model requests and GPU utilisation. That is a polite way of saying the bills currently arrive without a breakdown anyone can use.
Two structural details make the forecast worse than the headline rate suggests. Prepaid pools usually expire, so unused capacity is not a buffer. And overage enforcement is a cliff rather than a slope. Microsoft disables custom agents at 125% of prepaid capacity, with end users told the agent has reached its usage limit.
That second detail is the one buyers underprice. A budget overrun in seat software is an invoice conversation. A budget overrun in agent software is an outage. The survey of agent orchestration tooling shows how many meters a single workflow can now touch at once.
The evidence that opacity is now blocking sales
This is the part vendors should read twice. Confusing pricing is no longer a post-sale support cost. It is a pre-sale conversion cost.
Levelpath surveyed 300 US procurement and supply chain leaders at organisations actively buying AI software in June 2026. AI ranked as a top buying priority and as the slowest thing they buy. The most common AI purchase cycle was 16 to 20 weeks, against 7 to 10 weeks for standard software over $10,000.
The stakeholder count moved with it. 58% of AI purchases involve seven or more people, against 43% for standard software, and 28% involve eleven or more. Every additional approver is another person who has to be taught what a credit is.
Finance is in the room, and a token budget makes a veto more likely
G2 surveyed 1,038 B2B decision-makers in June 2026. Finance involvement in software buying rose from 31% to 46% in a single year, while information security involvement fell from 32% to 25%. The gate moved from whether the product is safe to whether the bill is survivable.
Then the counterintuitive finding. Among organisations that had set up a dedicated token or LLM budget, 54% reported a CFO blocking an approved purchase. Among those without one, the figure was 29%.
Read that carefully before drawing the obvious conclusion. Measuring token spend does not cause vetoes. It reveals the variance that was always there, in organisations already spending enough to need a budget line. The lesson for a vendor is not that transparency is dangerous. It is that the moment a buyer can see the variance, an unexplained rate card stops being tolerable.
The consequences are already priced in elsewhere in the survey data. 57% of procurement leaders reported spend problems with an AI vendor, 35% received higher-than-budgeted bills, and 9% terminated a vendor over price increases. Only 16% said they felt very or extremely confident managing AI costs.
Where this argument is weakest
Three objections, and the first is stronger than I would like.
The case for bespoke units is real
An AI vendor genuinely does have variable cost of goods. Seat pricing forces the vendor to absorb usage risk it cannot control, which either raises the floor price for light users or bankrupts the vendor on heavy ones. A credit is an honest attempt to pass through a cost that actually varies.
Buyers have noticed. G2 found that preference for outcome or usage-based pricing doubled from 11% to 23% in one year. Nearly half, 49%, had already been offered variable-cost alternatives. 52% said variable pricing improved their view of the vendor. The market is not asking for seats back. It is asking for units it can model.
So the complaint in this post is narrower than "consumption pricing is bad". It is that vendors invented incompatible units and then declined to publish the distribution of how many get consumed.
The second objection is about attribution. Not all of the extra ten weeks is pricing. Security reviews outrank contract negotiation as a reported cause of delay, at 58% against 52%, and data residency questions are genuinely new work. Pricing complexity is a substantial share of the gap, not the whole gap.
The third is about the evidence itself. Both buyer-side surveys here are vendor-sponsored research from companies that sell into this problem. Levelpath sells procurement software and G2 sells to software marketers. Sample sizes and fielding dates are disclosed, which is more than most, and the findings point the same way as the FinOps Foundation's independent survey. Treat the exact percentages as directional and the direction as well supported.
What a vendor should publish tomorrow
If you sell an AI product, the fastest available win is not a price cut. It is a denominator.
Publish the full list of billable events and their rates, as Microsoft does. Then publish the thing almost nobody does: the observed distribution of units consumed per interaction across live deployments, with sample size and time window. Not the average. The average is the least useful number in a skewed distribution, and consumption distributions are always skewed.
A buyer who can see that the tenth percentile fires 3 actions and the ninetieth fires 40 can build a scenario. A buyer given a single average builds a worst case instead, and then asks their CFO to approve it. That is how a deal reaches week 16.
Three more things cost nothing and remove weeks. Say whether unused units expire. Say what happens at the overage threshold, in operational terms rather than commercial ones. And offer a spend cap that throttles rather than disables, because the current alternative asks a buyer to choose between an unbounded bill and an unplanned outage.
The vendors that do this will win deals against better products. That is the uncomfortable implication of finance sitting in 46% of buying decisions. The same dynamic runs through the piece on seat compression rewriting SaaS pricing and the assessment of Salesforce in the agent era.
Frequently asked questions
Why is AI billing so hard to forecast?
Because the unit changes by vendor and the consumption changes by design. A seat is a countable thing you already track in a payroll system. A credit is a private currency whose drawdown depends on how an agent was built, which model it calls, and how many steps it takes. Microsoft's own rate card runs from 1 credit to 100 for a single response. Ensono's CFO put it plainly in August 2026: there is no perfect way to forecast AI spend.
What is the difference between credits, tokens and outcomes in AI pricing?
A token is a unit of model input or output, priced by the model provider. A credit is a vendor currency that sits on top of tokens and bundles compute, orchestration and margin into one number. An outcome is a business event, such as a resolved support conversation, that the vendor agrees to bill on. Tokens are comparable across providers. Credits and outcomes are not, because each vendor defines its own.
How much does a Copilot Credit cost?
Microsoft sells Copilot Credits in prepaid packs of 25,000 for $200 a month, which works out at $0.008 a credit. The pack price is the easy part. What a credit buys varies: a classic answer costs 1 credit, a generative answer 2, an agent action 5, tenant graph grounding 10, and a premium text and generative AI tool 100 credits per 10 responses. Unused credits do not roll over.
What counts as a resolution in AI agent pricing?
It depends on the vendor, and the difference is worth real money. Zendesk bills an automated resolution when an AI agent handles an interaction with no human involvement and the ticket stays closed through a 72 hour quiet period. Intercom bills an outcome when the customer confirms resolution, stops asking, or a workflow completes, including handoffs. Same word, different trigger, and the handoff clause moves the boundary.
Does outcome-based AI pricing remove budget risk?
It moves the risk, it does not remove it. Paying per resolved conversation means you stop paying for failures, which is a genuine improvement on paying per session. You still cannot forecast the bill, because your volume of resolvable conversations is not fixed and the vendor controls the definition of success. Outcome pricing converts a quality problem into a volume problem. Finance can budget for volume more easily, but not precisely.
What should procurement ask an AI vendor about pricing?
Four questions, in writing. What is the exact list of billable events and their rates? What is the observed distribution of units per interaction in comparable deployments, not the average? Do unused units roll over, and what happens at the overage threshold? Can we cap spend at the account level without disabling the service? Only 16% of procurement leaders negotiated cost caps in 2026, which is the cheapest of the four to fix.
How to price-proof the next contract
Two moves, both available before your next renewal call.
Build your own translation layer, because no vendor will build it for you. Take your three largest AI line items and restate each as cost per completed business event: cost per resolved ticket, per drafted document, per qualified lead. That number is comparable across vendors even though their units are not. It is the only column that belongs in a board pack, and it is the discipline applied in the analysis of where measurable AI return has shown up.
Then send one email to each vendor asking for the consumption distribution behind their quote, with sample size and period. The answer, or the absence of one, tells you more about your renewal risk than the rate card does. A vendor that has measured its own product can produce it in a day.
Related on this site
Metered agents are reshaping the software stack, not just the invoice. See what happens to point tools in the analysis of point SaaS absorption, and what unmanaged tool sprawl costs in the piece on 291 apps per company.
References
- Microsoft Learn, Copilot Studio billing rates and management, updated 3 August 2026. Used for every Copilot Credit rate, the reasoning model double meter, the 7,200 credit example and the 125% enforcement threshold.
- Microsoft, Microsoft 365 Copilot pricing, Copilot Studio, read 20 August 2026. Used for the $200 per 25,000 credit pack.
- Salesforce, Flexible pricing is a full circle moment for Salesforce, 23 May 2025. Used for the $0.10 per action Flex Credit rate.
- Intercom, Intercom pricing, read 20 August 2026. Used for the $0.99 per outcome rate and the definition of an outcome.
- Zendesk, Understanding outcome-based pricing, updated 12 May 2026. Used for the $1.50 automated resolution rate and the 72 hour quiet period.
- Levelpath, AI tops enterprise buying priorities yet takes the longest to buy, 9 July 2026. Survey of 300 US procurement and supply chain leaders, fielded June 2026. Used for cycle lengths, stakeholder counts, delay causes and spend-control figures.
- G2, 2026 Buyer Behavior Report, June 2026. 1,038 B2B decision-makers plus 55 marketer interviews. Used for finance involvement, CFO veto rates, the token budget split and pricing preference shifts.
- FinOps Foundation, State of FinOps 2026. 1,192 respondents representing more than $83 billion in annual cloud spend. Used for AI spend management adoption and the ranked challenges.
The weakest part of this source base: the two buyer-side surveys are vendor-sponsored research from firms that sell into this problem, and both report most-common response bands rather than medians, so the cycle lengths are modal not average. Every vendor rate quoted is a published list price. Contracted rates, volume tiers and overage multipliers are negotiated and almost never public. The Zendesk and Intercom resolution definitions come from each vendor's own marketing pages, which is the only place either publishes them.
Related reading