From Aryan Vatsa | Product & Market Analysis
Outcome-Based Pricing Sounds Great Until You Dispute the Outcome
On this page
Outcome pricing moves the risk of failure onto the vendor. It also hands the vendor the scoreboard. The outcome based pricing risk that matters is not the rate, it is the missing dispute mechanism. At Fin, an unanswered question becomes a billable outcome after 24 hours of customer silence, and the only record of that silence sits in the vendor's own telemetry.
Key takeaways
- The definition of the outcome is the price. Fin counts an assumed resolution when a customer disengages for 24 hours after its last answer, so a buyer who negotiates the rate and accepts the definition has negotiated the wrong term.
- Vendor telemetry is only presumptive evidence if the contract says so. Absent a named source of truth, an export right and a dispute window, the meter that generates the invoice is also the sole record of whether the invoice is correct.
- A miscount is a billing error, not a service level miss. Buyers who route counting disputes through the service level regime cap their own remedy at service credits, which is the wrong answer for being charged for work that did not happen.
- Your vendor's auditor now needs the same definition you do. Deloitte's technology spotlight of 4 June 2026 makes the definition of a successful outcome a revenue recognition judgment, which gives buyers unusual leverage to demand precision.
Written for the head of delivery or operations who signs the vendor paper and then owns the invoice every month. The core of this post is a clause list you can hand to legal.
What you actually agree to when you pay per outcome
Per seat pricing sells access. Outcome pricing sells a result, and a result has to be recognised by somebody before an invoice can be raised.
That recognition step is genuinely new. Nobody has to certify that a licence was used, so a seat contract never needed a definition of success.
Somebody has to decide, thousands of times a month, whether the thing you are paying for actually happened.
In nearly every agreement on the market today, that somebody is the vendor's own software. This is the same shift described in the analysis of seat compression in SaaS pricing.
The announcements price the outcome without defining it
HubSpot moved Breeze Customer Agent to $0.50 per resolved conversation, effective 14 April 2026, down from $1.00 per conversation. Chief customer officer Jon Dick framed it plainly in the company announcement: "Outcome-based pricing removes that risk. You pay when it works, full stop."
That page does not define what works means. Neither does MarTech's coverage of the change on 2 April 2026, which reports the price and the rationale without the counting rule.
I do not read that as evasion. A pricing announcement is not a contract, and the operative definition sits in help centre documentation and order forms. The term that determines your bill is never in the material you are shown while deciding.
Vendors publish one half of the unit and not the other
Zendesk's pricing page defines the unit precisely and then stops. It states that "paying per automated resolution means that you pay only for customer requests that were successfully resolved by the AI agent, without any escalation to a human agent". It does not publish a per resolution rate or an overage rate anywhere on the page.
Fin does the reverse and publishes both, listing four billable outcome types with prices attached. The pattern across the three vendors is inconsistent rather than dishonest, and inconsistency is what defeats a comparison.
The vendor's telemetry is the only record, and that is the real exposure
The commercial pitch for outcome pricing is risk transfer. You pay for results, so a failed agent costs you nothing.
The pitch holds only if failure is observable to both parties. When the vendor's platform is the sole instrument of measurement, risk has not transferred so much as changed shape.
The meter ships inside the product
An outcome meter is a piece of the vendor's application. It runs on the vendor's infrastructure, applies the vendor's classification logic, and writes to the vendor's data store.
None of that is improper. It is the same party grading the exam and printing the certificate, which is a structure procurement teams would reject instantly in any other spend category.
Mayer Brown's February 2026 note on contracting for agentic AI solutions makes the same observation from the legal side. It recommends that providers maintain decision logs for AI agent actions in structured, legible form, so the buyer can audit decisions rather than accept a summary.
Silence is the weakest evidence a meter can accept
Fin's published outcome documentation is unusually clear. A resolution is counted when the customer either confirms the answer helped, or exits the conversation without asking for more help.
The second case is the assumed resolution, triggered by 24 hours of disengagement. A customer who read an unhelpful answer, gave up and solved the problem elsewhere produces exactly the same telemetry as a customer who was helped.
Fin does provide a reversal. Its documentation states that if a resolved conversation is later reopened by the customer seeking further assistance, even across billing periods, that resolution is deducted and not charged.
That is a good clause, worth asking every vendor to match. It also only catches the customers who come back, and those are the minority in most support datasets.
A model update can move the count without touching the contract
Outcome counts depend on model behaviour, prompt content and guardrail configuration. All three change on the vendor's release schedule, not yours.
Mayer Brown's June 2026 note on agentic AI implementation deals states the underlying problem directly. "Because agentic AI outputs are non-deterministic, a single successful test is not enough to demonstrate that a solution will perform as required in production."
If a new model version becomes more willing to declare a conversation handled, your invoice rises without your volume rising, and nothing in a standard subscription contract requires the vendor to tell you.
Inference economics push in the same direction, because a vendor whose margin is squeezed has an incentive to raise counted outcomes per unit of compute. That pressure is covered in more depth in the piece on inference costs and AI margins.
Seven clauses that make a disputed outcome resolvable
The working list assumes a mid market deal where you have some negotiating room and no appetite for a six month legal exercise.
Take the first four if you can only win four. They convert an argument about judgement into an argument about evidence.
| Clause | What it has to say | What goes wrong without it |
|---|---|---|
| 1. Outcome definition and counting rules. | The billable event, named exclusions, one charge per unit of work, and the treatment of a reopened request. | Vendor product documentation becomes the contract, and it can be edited without your consent. |
| 2. Named source of truth. | Which system's record governs the invoice, and which record wins if two systems disagree. | Every dispute becomes a debate about whose dashboard is right, with no tie-break. |
| 3. Export and retention rights. | Event level records, decision traces, tool use logs and model version identifiers, in a structured format, on a stated cadence. | You see a total and never the line items behind it, so you cannot check the bill at all. |
| 4. Dispute window, hold back and clawback. | A stated number of days to challenge, the right to withhold only the disputed amount, and a vendor response deadline. | Disputing an invoice puts you in breach for non payment, which ends the conversation quickly. |
| 5. Sampling and verification. | The right to have a stated share of billed outcomes reviewed each period against your own records. | You can only challenge individual cases you happened to notice, never a systemic error. |
| 6. Change control on models and prompts. | Advance notice of material changes, version pinning where offered, a re-baselining trigger and rollback authority. | Your counted outcomes move for reasons that never appear anywhere in your own operations. |
| 7. Remedy that is not only a service credit. | Credits for performance misses, plus billing correction for miscounts and a termination right at a stated threshold. | Systemic overcounting is treated as a service level miss and refunded at credit rates, not cash. |
This checklist is drafting guidance assembled from published law firm commentary and vendor documentation. It is not legal advice and it is not a substitute for counsel who has read your order form.
Clauses 1 and 2: the definition and the tie-break
Start by writing the definition into the agreement itself rather than incorporating a help centre page by reference. A page you do not control is not a term you have agreed.
Then name the exclusions explicitly. Greetings, duplicate contacts from the same requester, tests, internal traffic and conversations that reach a human should all be listed. Each one is a category that reasonable people classify differently.
Clause 2 is the one buyers forget. Where your own service desk and the vendor's meter disagree, the contract should say which record governs and what happens when the gap exceeds a material threshold.
Clauses 3 and 4: evidence and a way to argue about it
Mayer Brown's June 2026 note lists what observability should cover for autonomous agents: agent logs, decision traces, tool use records, prompts and system instructions, model and version identifiers, evaluation results and drift indicators.
Ask for that list as an export, not as a dashboard. Agent telemetry that lives only in the vendor's console is a reporting feature, and what you need is evidence you can hold.
Clause 4 gives the evidence somewhere to go. A 30 day dispute window, a 10 business day vendor response deadline and a right to withhold only the contested amount are unremarkable terms in any other services agreement. Here they are frequently absent.
Clause 7: keep miscounting out of the service level regime
This is the drafting error I see most often, and it is expensive. Buyers negotiate hard on service levels, win service credits, and then discover that credits are the exclusive remedy for everything including a billing error.
Mayer Brown notes that providers may accept service credits as an exclusive remedy where outcome metrics are clearly defined. That is a reasonable trade for performance misses. It is not a reasonable trade for being charged for work that did not occur.
Separate the two regimes. Performance failures earn credits, and miscounts get corrected in cash, because a credit against future spend is worth less than the money and it quietly rewards the party that made the error.
Weak wording and wording that works
Standard vendor paper is not malicious, it is simply drafted by the party with better information. These four substitutions carry most of the value.
| Common vendor wording | Wording that survives a dispute |
|---|---|
| "Provider's usage reports shall be conclusive evidence of the outcomes delivered". | "Provider's reports are presumptive evidence. Customer may dispute any line within 30 days, and Provider shall supply the underlying event records within 10 business days". |
| "A Resolution means a conversation that concludes without escalation to a human agent". | "A Resolution means [definition], excluding [listed categories], counted once per requester per issue, and reversed if the same requester reopens the same issue within [N] days". |
| "Provider may modify the Services, including models and prompts, at any time". | "Provider shall give [N] days notice of changes materially affecting outcome measurement. Customer may re-baseline within one billing period following such a change". |
| "Service credits are Customer's sole and exclusive remedy". | "Service credits are the sole remedy for service level failures. Billing errors, including miscounted outcomes, are corrected by invoice adjustment or refund". |
Bracketed values are deliberately blank. The numbers that fit a 2,000 conversation a month deployment do not fit a 200,000 one.
Outcome SLAs are not uptime SLAs
Availability is the wrong measure for an agent. A platform can be up for the whole month and still classify a quarter of its outcomes wrongly.
Mayer Brown's February 2026 note proposes operational measures instead. Its worked examples include 99% of invoices processed correctly against the purchase order, 99% of support tickets actioned within the required service window, and fewer than 1% of autonomous actions leading to consumer complaints.
Its June note goes further and separates performance into four layers. Splitting them matters, because a vendor who commits only to the platform layer has committed to almost nothing you care about.
| Layer | What it measures | Typical coverage today |
|---|---|---|
| Implementation. | Deliverables, milestones, acceptance criteria, deployment dates. | Covered in statements of work, rarely tied to the outcome meter. |
| Operational. | Support, uptime, latency, incident response, drift management. | Almost always covered, and the least useful of the four. |
| Agent performance. | Completion rate, human handoff rate, rework rate, tool use reliability. | Sometimes reported, seldom committed to as a service level. |
| Business outcome. | Cycle time, cost reduction, error reduction, customer experience. | Quoted in the sales deck, rarely in the agreement itself. |
Acceptance testing needs the same treatment. Mayer Brown recommends testing against an agreed evaluation set with defined thresholds for accuracy and latency, assessed over time in production rather than once in a demo environment.
That is the discipline missing from most agent purchases, and the same measurement gap described in the piece on where measurable GenAI return has actually shown up. If you cannot state the threshold, you cannot claim the miss.
Why the vendor's auditor now needs your definition too
Here is the piece of leverage most buyers do not know they have. The definition of a successful outcome is no longer only a commercial term, it is a revenue recognition input.
Deloitte published a technology spotlight on accounting for outcome-based pricing in an agentic AI software product on 4 June 2026. It describes outcome-based pricing as arrangements where payment is tied to achieving contractually defined successful outcomes rather than to providing access to the agent.
The document identifies the key judgment as whether the vendor's promise is a stand ready obligation to provide continuous access, or an obligation to deliver a specified quantity of successful outcomes. The first recognises revenue over time, the second as outcomes occur.
It then states the requirement that matters to you: the contract must define what constitutes a successful outcome with sufficient specificity to allow both parties to determine when the obligation is satisfied.
Use that. A vendor resisting a precise definition is resisting something its own auditors will ask for anyway, and saying so in a negotiation is more effective than arguing about fairness.
Outcome revenue is variable consideration subject to the constraint, so vendors have an accounting reason to prefer definitions that produce predictable counts, and predictable is not the same as accurate. The platform incentives around this are set out in the analysis of Salesforce in the agent era.
Where this argument is weakest
Four honest problems with everything above.
There is no public data on how often outcome invoices are actually disputed. My case rests on documented definitions, published law firm guidance and the structure of the incentive. It does not rest on a measured dispute rate, because none exists that I could verify, and a structural risk is not a realised loss.
Outcome pricing is still usually better for the buyer than seats. Paying only for work that the vendor believes happened is a better starting position than paying for licences nobody opens. The dispute problem is a second order issue inside a first order improvement.
Vendors have legitimate reasons to resist full trace access. Prompts and system instructions are commercially sensitive, decision traces can contain personal data, and a raw export creates a privacy obligation you may not want either. Ask for structured event records and model version identifiers first, and treat full prompt disclosure as a separate negotiation.
Most buyers will never get this paper. Below a certain deal size you sign the standard terms, and a clause list is aspirational. The EU AI Act logging duties in Articles 12 and 19 apply only to high-risk systems under the Act, and a support agent generally is not one. Those obligations are a useful template rather than a right you can claim by default.
The honest position is that this is drafting hygiene with an uncertain payoff, in a market where the same law firm guidance concedes that market standards have not yet developed. I would still take the four hours to negotiate clauses 1 to 4.
Frequently asked questions
What is outcome-based pricing for AI agents?
Outcome-based pricing charges for a defined result rather than for access or consumption. Deloitte describes it as payment tied to achieving contractually defined successful outcomes rather than to providing access to the agent. Current examples include HubSpot Breeze at $0.50 per resolved conversation and Fin at $0.99 per resolution. The commercial risk sits with the vendor, and the measurement usually sits with the vendor as well.
What counts as a resolution in an AI agent contract?
It depends entirely on the vendor, which is the problem. Fin counts a resolution when the customer confirms the answer helped, or when the customer disengages for 24 hours after Fin's last answer without asking for more help. Zendesk defines an automated resolution as a request resolved without escalation to a human agent. Neither definition tests whether the customer's problem was actually solved.
How do you dispute an AI agent invoice?
You need a clause, because the default position is weak. Ask for a stated dispute window of around 30 days and a vendor obligation to supply the underlying event records within a fixed number of business days. Also ask for the right to withhold only the disputed amount without triggering a payment breach. Without a hold back right, disputing an invoice puts you in breach and the argument ends there.
Should service credits be the only remedy in an AI agent SLA?
No. Service credits are a reasonable exclusive remedy for performance failures where the metric is clearly defined. They are the wrong remedy for a miscounted outcome, which is a billing error and should be corrected by invoice adjustment or refund. Keep the two regimes separate in the drafting, because a single exclusive remedy clause otherwise converts every overcharge into a discount on future spend.
What audit rights should a buyer ask for in an agentic AI contract?
Mayer Brown's 2026 guidance lists agent logs, decision traces, tool use records, prompts and system instructions, model and version identifiers, evaluation results and drift indicators, with access continuing through termination plus a tail period. Ask for these as a structured export on a stated cadence rather than as dashboard access, and add a sampling right so you can test a share of billed outcomes each period.
Is outcome-based pricing cheaper than per-seat pricing?
Often yes, and the comparison depends on volume rather than on the model. Outcome pricing removes the cost of unused licences and adds a cost that scales with activity, so a high volume support operation can pay more than it did on seats. Model both curves against your own ticket volume before signing, and check whether the vendor's committed volume rate is materially below its pay as you go rate.
Where to start this week
Two tasks, both short, both doable before your next renewal conversation.
First, find the operative definition for every outcome you are already paying for. It will be in a help centre article rather than in your contract, so save a dated copy of that page. It is the only version you can prove you agreed to.
Second, run one month of your own records against the vendor's count. Pick the single outcome type with the highest volume, compare the vendor total to your service desk total, and write down the gap. If the gap is under 2% you have a reporting question, and if it is over 10% you have a negotiation.
The same exercise is worth running against any orchestration layer you have deployed, which is the subject of the review of agent orchestration and automation tools.
The one sentence to put in the order form.
Provider's usage reports are presumptive rather than conclusive, Customer may dispute any line within 30 days, and Provider shall supply the underlying event records within 10 business days.
References
- Intercom Help Centre, Fin AI Agent outcomes, accessed 20 August 2026. Used for the resolution definition, the 24 hour assumed resolution window, the reversal on reopen and the $0.99 to $9.99 outcome prices.
- HubSpot, HubSpot's Customer Agent and Prospecting Agent: now you pay when the task is complete. Used for the $0.50 per resolved conversation price, the 14 April 2026 effective date and the Jon Dick quote.
- MarTech, HubSpot moves to outcome-based pricing for some Breeze AI agents, 2 April 2026. Used to confirm the pricing change and the absence of a published counting rule in the coverage.
- Zendesk, Pricing, accessed 20 August 2026. Used for the automated resolution definition and for the absence of a published per resolution rate.
- Mayer Brown, Contracting for Agentic AI Solutions: Shifting the Model from SaaS to Services, 17 February 2026. Used for decision log guidance, the worked service level examples and the service credit position.
- Mayer Brown, Key Contract Issues in Agentic AI Implementation and Integration Deals, 16 June 2026. Used for the four commitment layers, the observability list, acceptance testing guidance and the non-determinism quote.
- Deloitte, Technology Spotlight: Accounting for Outcome-Based Pricing in an Agentic AI Software Product, 4 June 2026. Used for the definition of outcome-based pricing, the stand ready judgment and the specificity requirement.
- EU AI Act, Article 19, Automatically generated logs. Used for the six month minimum retention period for high-risk AI system logs.
The weakest thing about this source base is that it contains no dispute data. Two of the eight sources are law firm commentary describing what contracts should say rather than what signed contracts do say, and the three vendor pages are list pricing rather than negotiated terms. The 0 of 3 figure in the stat strip is my own count across three public pages on 20 August 2026. That is a very small sample and should be read as an illustration rather than as a market finding.
Related reading