From Aryan Vatsa | Product & Market Analysis

AI Agent Pricing for a Product Nobody Trusts Yet: The Four-Rung Ladder

On this page

Only 46% of people say they are willing to trust AI systems, and Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027. That is a pricing problem before it is an engineering problem. AI agent pricing in a low-trust category is not a rate card, it is a schedule for moving risk off the buyer and onto you.

Key takeaways

  • Trust, not value, sets the ceiling on your first price. Only 46% of people say they are willing to trust AI systems, in a study of more than 48,000 people across 47 countries run by KPMG and the University of Melbourne.
  • Pure outcome pricing asks the buyer for two acts of faith at once. They must trust the agent and trust your attribution. That is why the model works in ticket deflection, where the counterfactual is already metered, and stalls where it is not.
  • The free pilot is the most expensive price you can set. It signals the work is worth nothing and does not prevent the cancellation anyway. Gartner expects over 40% of agentic AI projects to be cancelled by the end of 2027 on cost, unclear value and weak risk controls.
  • Put the human reviewer inside the fee rather than outside it. The strongest agent in Carnegie Mellon's benchmark finished 30% of consequential office tasks without help. Supervision is part of the product, and buyers pay for supervision they can see.
46%Willing to trust AI systems, against 66% who use it regularly. Source: KPMG and University of Melbourne, April 2025.
40%+Agentic AI projects Gartner expects to be cancelled by end 2027. Source: Gartner, June 2025.
30%Consequential office tasks the strongest agent finished autonomously. Source: TheAgentCompany, 2025.

What low trust actually costs you at the negotiating table

Short answer: in a low-trust category you are not pricing software, you are pricing risk transfer. Charge a real fee for a short, scoped pilot, keep a human review budget inside the price, and move to per-outcome billing only once the outcome has an agreed meter. Trust buys the right to charge for results.

Most pricing advice for AI founders starts from value. Work out what the agent saves, take a slice, defend it in the deal. That assumes the buyer believes your number, and in this category they do not. The distance between what your agent is worth and what a buyer will underwrite is the real constraint on your price.

That distance is not irrational. A buyer signing a large agent contract is putting their own judgement on the line inside their own company. The software can work perfectly and the project can still be cancelled, because value was never the thing being tested.

The discount you are already giving

Every low-trust deal contains a discount you did not decide to give. It shows up as a shorter term, a lower committed volume, a pilot that never ends, or a clause letting the buyer walk after 90 days without cause.

Those concessions cost more than a headline discount, because they are invisible in your average selling price and lethal in your net revenue retention. A 30% price cut is a number you can see. A twelve-month contract that is really four one-month contracts is not.

A free pilot is a price, and it is the worst one

The free pilot is the single worst pricing decision I see AI founders make. It is almost always framed as a growth tactic. It looks like removing friction. It is setting your price to zero at the exact moment the buyer is forming their view of what the work is worth.

Free pilots also select for the wrong buyer. The people who take them are disproportionately the ones with no budget, no mandate and no baseline data, which is why so many convert into nothing. A paid pilot filters for a sponsor who has already moved money.

The evidence that trust is the binding constraint

46% is the ceiling on unsupervised adoption

The KPMG and University of Melbourne study surveyed more than 48,000 people across 47 countries. It found that 66% use AI regularly while only 46% are willing to trust it. Those findings are not in conflict. They describe a population that uses the tool and supervises it.

The same study found 66% of employees rely on AI output without checking accuracy, and 56% report making mistakes at work because of AI. Read those together and you get the shape of the risk your buyer is pricing. Not that the agent fails, but that nobody notices.

This is a general-population survey rather than a survey of enterprise buyers, so treat it as directional for procurement behaviour rather than proof of it.

30% autonomy is the engineering reality

Researchers at Carnegie Mellon and collaborating institutions built TheAgentCompany, a benchmark of 175 long-horizon professional tasks inside a simulated software company. The tasks span engineering, project management, data science, administration, HR and finance. The strongest agent tested completed 30% of those tasks autonomously.

Simpler tasks were solved and difficult long-horizon work was not. It is a simulated environment, so it is not a claim about any specific product, and narrow production deployments do better than the general case.

It still sets the frame. If roughly two thirds of consequential multi-step work needs a person somewhere in the loop, a pricing model that assumes full autonomy is pricing a product that does not exist yet.

Three numbers that set the price before you do Different sources, different methods. Read them as a direction, not one measurement. 46% Willing to trust AI systems KPMG and Univ. of Melbourne, 2025 30% Office tasks done autonomously TheAgentCompany benchmark, 2025 40% Agent projects forecast cancelled Gartner forecast to end 2027
The third dial is a forecast, not a measurement. It is the one most often quoted as fact. Read it as Gartner's view of where risk controls are weakest.

Why pure outcome pricing is premature in most categories

Outcome pricing is the fashionable answer and a genuinely good idea in the right category. Bret Taylor of Sierra makes the cleanest version of the case in a March 2026 interview. Where an outcome is measurable, he argues, it aligns the vendor with the customer's business. A $10 or $20 phone call can become a 10 or 20 cent interaction.

I disagree with the common extension of that argument, which is that outcome pricing is the natural end state for every agent category. It is the end state for categories where the outcome already has a meter attached, and that is a much smaller set than the discourse implies.

Outcome pricing asks for two acts of trust at once

When you bill per outcome, the buyer has to accept two things. First, that the agent produced the result. Second, that the result would not have happened anyway.

The second is the hard one. In support, the counterfactual is a ticket a person would otherwise have handled, and the cost of that person is already in a spreadsheet. In sales, marketing, finance operations or engineering, the counterfactual is contested, and every invoice becomes an attribution argument.

I would rather lose a deal at a modest base fee than win it on outcomes and spend the following year arguing about which of us caused the number. Attribution disputes do not only cost margin, they consume the relationship that renewal depends on.

Where per-outcome pricing already works

Two vendors publish rates, which makes them the only honest reference points in a market full of unpublished numbers. Intercom prices its Fin agent at $0.99 per resolution, with no setup, integration or platform fee. Salesforce prices Agentforce at $2.00 per customer-facing conversation, or about $0.10 per standard action through Flex Credits.

Notice what those units have in common. A resolution and a conversation are both events the buyer's existing systems already count. Neither vendor had to invent a meter, and inventing one is the part founders underestimate.

Published AI agent price points, August 2026
Vendor and unitPriceNote
Intercom Fin, per resolution$0.99No setup or platform fee. 50-outcome monthly minimum.
Intercom Fin, per qualification$9.99A different outcome type at 10 times the resolution rate.
Agentforce, per conversation$2.00Customer-facing agents only, consumption model.
Agentforce, per standard actionAbout $0.1020 Flex Credits at $500 per 100,000 credits.
Agentforce, per user add-on$125 per monthUnmetered employee agent usage, sold alongside the meter.

Prices are from the vendors' own pricing pages, retrieved 20 August 2026. Sierra is deliberately excluded: its per-resolution figure circulates in secondary coverage and the company publishes no rate, so there is nothing to verify. Selling a per-user add-on and a consumption meter side by side is covered in the piece on what happens to seat-based pricing when agents do the work.

Published prices per unit of agent work Log scale. Each gridline is a tenfold step. The spread runs from cents to dollars. $0.10 $1.00 $10.00 Agentforce action $0.10 Fin resolution $0.99 Agentforce conversati… $2.00 Fin qualification $9.99 Source: fin.ai and salesforce.com pricing pages, retrieved 20 August 2026.
The tenfold gap between a Fin resolution and a Fin qualification is the point. Both are outcomes. One is worth more to the buyer, and the price says so.

The four-rung pricing ladder

Here is the structure I would use. Each rung moves risk from the buyer to you, and each rung is paid for by a higher effective rate. You do not skip rungs, and you do not stay on one forever.

A staged pricing ladder for a low-trust agent category
RungWhat you chargeWhat the buyer risksWhat you prove to climb
1. Paid diagnostic pilotFixed fee, 4 to 8 weeks, one workflow, credited against year oneA small budget line and staff timeA recorded baseline and a measured task success rate
2. Capped subscriptionPlatform fee including a stated volume of supervised actionsA year of budget, no headcount decisionReview load falling across three straight months
3. Base plus per outcomeSmaller base fee, per-outcome charge above an agreed floorVariable spend they forecast but cannot fixA meter both sides accept, plus clean escalation data
4. Outcome onlyPer resolved outcome, with a minimum commitmentVery little, which is the pointAttribution that survives their internal audit

This ladder is a proposed structure, not a measured finding. It is built from the published pricing models above and from the contracting practice Mayer Brown described in February 2026. Treat the rung boundaries as a starting point to argue with.

Each rung buys a higher rate by taking risk off the buyer Illustrative structure. Bar height shows how much delivery risk the vendor carries. Rung 1 Paid pilot Fixed fee, 4 to 8 weeks. Rung 2 Capped subscription Review budget included. Rung 3 Base plus outcome Metered above a floor. Rung 4 Outcome only Per resolved outcome. BUYER CARRIES THE RISK YOU CARRY THE RISK
The mistake is jumping from rung one to rung four because a competitor advertises outcome pricing. You cannot price attribution you cannot yet produce.

Pricing rung one, where most of the damage is done

The pilot is the rung founders get wrong most often, and the cheapest one to fix. Three decisions matter.

First, the fee. Price it high enough that a director has to think, and low enough that it does not need a committee. In practice that is usually one to three months of the annual contract value you eventually want, structured as a fixed fee and credited in full against year one on conversion.

Second, the scope. One workflow, one team, one success metric, with a written definition of what counts as the agent doing the work. Vague scope is how an eight-week engagement becomes a nine-month unpaid deployment.

Third, the deliverable. A pilot that ends in a demo has produced nothing you can sell against. A pilot that ends with a recorded baseline, a measured task success rate on the buyer's own data and a count of escalations has produced the evidence for rungs two and three.

Be stubborn about that last point. The pilot's real product is the measurement, and the measurement is what lets you charge more later. If the buyer's own numbers are missing, price the baseline work explicitly, because building the counterfactual is real consulting. The same discipline applies when you are competing against a team's instinct to build it themselves, covered in the analysis of build versus buy for coding agents.

Price the human reviewer in, not out

The instinct is to hide the human. Reviewers look like a confession that the agent is not ready, so founders bury the cost in gross margin and hope it shrinks. That is backwards at low trust.

If I were pricing a new agent today, I would put supervision on the invoice as a named line item and defend it. A buyer looking at 46% public trust and a 40% cancellation forecast is not reassured by the absence of a reviewer. They are reassured by a stated review rate, a stated escalation policy, and a price that visibly includes both.

There is a commercial argument as well as a psychological one. A review line is a margin lever you control. As the agent improves you cut the review rate, hold the price, and the improvement lands in your gross margin instead of in a discount. Hide the review and you have nothing to give back later except money.

The regulatory clock does not rescue you. The EU's Digital Omnibus on AI, in force from 27 July 2026, deferred the high-risk obligations that include human oversight. Stand-alone Annex III systems now have until 2 December 2027, and AI embedded in regulated products until 2 August 2028. Deferred is not cancelled, and enterprise buyers are already writing oversight requirements into contracts ahead of the deadline.

The contract does half of the pricing work

Pricing conversations at low trust fail on terms more often than on rates. Mayer Brown's February 2026 analysis argues that agentic AI contracts are moving away from the software-as-a-service template toward outsourcing-style structures, which is the right diagnosis of what buyers are asking for.

The specifics matter for pricing. The firm proposes defining the service as tasks completed rather than software provided, and replacing as-is disclaimers with a warranty of professional, workmanlike performance. It also proposes shifting service levels from uptime to accuracy and timeliness, and giving the buyer access to decision logs showing why the agent chose what it chose. Read that list as a price sheet, not a legal checklist.

Service credits are the sharpest item on it. A credit tied to an accuracy or escalation threshold is risk sharing you can quantify, cap and price. A discount is risk sharing you never get back. I would trade a 10% headline discount for a service credit regime every time, because the credit only pays out when I have underdelivered and the discount pays out always.

Two boundaries deserve careful drafting. Delegation of authority defines what the agent may do without a person, and policy guardrails define what it may never do. Both are pricing boundaries as much as safety ones, because the wider the delegation, the more of the buyer's risk you have absorbed and the more the work is worth.

Where this argument is weakest

The case for skipping straight to outcomes

If your category already has a meter, the ladder wastes time and a competitor selling on outcomes takes your deals. Support is the clearest case. The cost of a human-handled contact already sits in the buyer's budget, the counterfactual is uncontested, and a per-resolution price is easier to approve than a subscription because it needs no capacity forecast.

In that world, staging is not prudence, it is slowness. The published rates from Intercom and Salesforce exist because those categories cleared the attribution problem years ago through call centre economics.

What would change my mind

Three things, all observable. If agent completion on long-horizon professional benchmarks moves from roughly 30% toward 70%, the supervision line stops being defensible and the ladder collapses into two rungs. If Gartner's cancellation forecast lands well under 40%, the trust discount was smaller than assumed and staged pricing was overcautious.

And if standard outcome meters emerge for functions beyond support, attribution stops being the bottleneck and rung three becomes the entry point rather than the destination. My evidence base here is two surveys, one benchmark, two published price lists and a law firm's reading of live deals. That is enough to build a structure on, and not enough to call it proven.

Frequently asked questions

How should I price an AI agent that customers do not trust yet?

Price the risk transfer, not the software. Start with a paid, scoped pilot that produces a baseline and a measured task success rate. Then sell a capped subscription that includes a stated volume of supervised actions. Move to per-outcome billing only once the outcome has an agreed meter.

Is outcome-based pricing better than usage-based pricing for AI agents?

Neither is better in the abstract. Outcome pricing aligns your revenue with the buyer's result and works where the outcome is already metered, as in support resolutions. It also asks the buyer to trust the agent and to trust your attribution, which is two leaps at once. In a low-trust category, meter only what both sides can count.

Should I offer a free pilot for my AI agent?

Rarely. A free pilot signals the work has no value, attracts buyers who were never going to pay, and gives you no evidence about willingness to pay. Charge a fixed fee for a scoped pilot instead, and credit it in full against the first year if the buyer converts.

How much do AI agents cost per outcome?

Published rates vary by unit. Intercom prices its Fin agent at $0.99 per resolution, with no setup or platform fee and a 50-outcome monthly minimum. Salesforce prices Agentforce at $2.00 per customer-facing conversation, or about $0.10 per standard action through Flex Credits. Many vendors publish no rate at all.

What is risk sharing in AI agent pricing?

Risk sharing means the vendor carries part of the cost when the agent underperforms. In practice it appears as service credits tied to accuracy and escalation reliability, and a fee that visibly includes human review. It also covers indemnities for breaches of the agent's delegated authority, and a pilot fee credited against year one.

Does the EU AI Act require human review of AI agents?

Not yet for most systems. The EU's Digital Omnibus on AI, in force from 27 July 2026, deferred the high-risk obligations that include human oversight. Stand-alone Annex III systems have until 2 December 2027, and AI embedded in regulated products until 2 August 2028. Build the review capability now anyway.

Where to start this week

Open your last five closed-lost and closed-won deals and write down the real commercial shape of each. Not the list price: the actual term length, the actual committed volume, and the actual exit clause. That is your current rung, and it is usually lower than founders think.

Then take your standard pilot and put a number on it. Pick the fee, define the single workflow, and write the one sentence saying what the pilot delivers on its final day. If that sentence describes a demo rather than a measurement, rewrite it before your next call.

Related on agent economics

Pricing is downstream of where value settles in the stack. See how agent ecosystems absorb point SaaS and why workflow moats outlast data moats.

References

  1. KPMG and University of Melbourne, Global study reveals trust of AI remains a critical challenge, April 2025. More than 48,000 people across 47 countries. Used for all trust figures.
  2. Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Poll of more than 3,400 organisations. Used for the cancellation forecast.
  3. Xu et al., TheAgentCompany: Benchmarking LLM Agents on Consequential Real World Tasks, arXiv 2412.14161, revised September 2025. Used for the 175-task count and the 30% figure.
  4. Intercom, Fin pricing, retrieved 20 August 2026. Used for the per-resolution and per-qualification rates and the minimum commitment.
  5. Salesforce, Agentforce pricing, retrieved 20 August 2026. Used for the per-conversation rate, Flex Credits arithmetic and the add-on price.
  6. Mayer Brown, Contracting for agentic AI solutions, 17 February 2026. Used for warranties, service levels and delegation of authority.
  7. Gibson Dunn, EU AI Act omnibus agreement: postponed high-risk deadlines, 2026. Used for the deferral dates under the Digital Omnibus on AI.
  8. A Cheeky Pint, Bret Taylor of Sierra on AI agents and outcome-based pricing, 10 March 2026. Used for the outcome-pricing argument and per-contact cost comparison.

The weakest part of this source base is the pricing comparison. Only Intercom and Salesforce publish rates, so the table describes two vendors rather than a market, and the per-contact cost figures attributed to Sierra are a vendor executive's characterisation rather than measured data.

AV
Aryan Vatsa
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading