From Aryan Vatsa | Product & Market Analysis
Post-SaaS Procurement: The Questions CIOs Now Ask Before Features
On this page
Software procurement in 2026 starts with a question that barely existed in 2023: what does this cost when usage triples? Redpoint's March 2026 survey of 141 CIOs found 45% of AI spend is drawn from existing software budgets rather than new ones. The buying process changed because the money changed.
Key takeaways
- AI budget is mostly not new budget. 45% of AI spend comes out of existing software lines and 54% of CIOs are consolidating vendors, per Redpoint's survey of 141 CIOs in March 2026.
- The unit of price moved from the seat to the action. Salesforce charges $0.10 per Agentforce action, replacing a flat $2 per conversation, which makes cost prediction harder than price negotiation.
- Data rights are the redline, not the footnote. The US federal AI acquisition memo requires contracts to state data ownership and IP rights explicitly. Most commercial master agreements still leave both silent.
- Exit cost is the strongest single disqualifier. A vendor who cannot export your data, prompts, configuration and audit trail in a documented format is a single point of failure whatever the demo showed.
What "post-SaaS procurement" actually means
Post-SaaS procurement evaluates software on three things before anyone opens the feature comparison. What it does autonomously. What it costs per unit of work. And what rights you keep over your own data. It is not a rejection of SaaS. It is a different order of questions, driven by the fact that AI money is taken from the software line rather than added to it.
The budget moved first and everything else followed. Redpoint's March 2026 survey of 141 CIOs is the clearest read on where the money comes from. 45% of AI spend is reallocated from existing software budgets, 54% of CIOs are consolidating vendors, and only 3% expect AI to increase their vendor count.
Read those three numbers together. AI is not arriving as a new category alongside the old stack. It is arriving as a claim on the same budget, held by the same person, in the same renewal cycle.
That changes the meeting. A renewal used to be an administrative event with a discount attached. It is now the moment a buyer decides whether a category still deserves a line item at all. That is the same pressure described in the analysis of seat compression in SaaS pricing.
The practical effect is that incumbents are re-competed by default. Buyers are not uniformly hostile to them. The same survey found 54% of CIOs would rather their existing vendor added AI than switch to an AI-native alternative. Familiarity still carries weight, and switching costs are real.
But the burden of proof inverted. The incumbent now has to show why the category still needs a separate vendor. That question is explored further in the piece on point tools being absorbed into agent ecosystems.
The questions that now come before features
Feature checklists have not disappeared. They have been demoted below a set of questions that were previously handled after signature, or not at all.
| What used to decide it | What decides it now | Where the old test still wins |
|---|---|---|
| Feature parity checklist | Inventory of actions the software takes without a human, and what happens when one fails | Regulated workflows, where the feature list is effectively the compliance list |
| Price per seat | Cost per unit of work at forecast volume, with a rate card | Stable headcount and bounded usage, where seats are simpler and cheaper to administer |
| Reference customers | Production evidence: latency under load, error rates, escalation paths | Implementation risk, which references genuinely predict better than benchmarks do |
| Security questionnaire | Data rights, training use, retention windows, output ownership | Baseline hygiene, which the questionnaire still catches efficiently |
| Product roadmap | Model dependency, version pinning and deprecation notice periods | Long platform bets, where direction matters more than this quarter's release |
Can it act, or only answer?
The first question is whether the software takes actions or produces text. The distinction is doing real work in evaluations because the failure modes are completely different.
An assistant that drafts a wrong reply wastes a minute. An agent that issues a wrong refund moves money. So the useful demand is not "show me the agent", it is "give me the list of actions it can take, the permissions each one needs, and the behaviour when the action fails".
Gartner has been blunt about the noise here. It describes "agent washing", the rebranding of assistants, chatbots and existing automation without substantial agentic capability. It predicts that more than 40% of agentic AI projects will be cancelled by the end of 2027. An action inventory is the cheapest available test for that, because a vendor without one cannot fake the answer in a meeting.
What does it cost at ten times the pilot?
The second question is volume. Consumption pricing makes the pilot invoice almost meaningless as a guide to the production invoice.
A pilot runs on a small team, a narrow workflow and careful prompts. Production runs on retries, long context, tool calls that fail and get repeated, and users who paste in whole documents. The unit price does not change. The number of units does.
This is why the rate card matters more than the discount. A 20% discount on an unknown quantity is not a commercial term, it is a decoration.
The third question is rights, and it is the one most often deferred. Two defaults are worth checking on every agreement: whether your prompts and outputs train the vendor's models, and who owns the output. Enterprise tiers commonly disable training on customer data. But the setting is often a configuration rather than a contractual promise. Those are not the same thing when a vendor is acquired or changes its terms.
Consumption forecasting is the skill nobody hired for
Procurement teams are good at negotiating rates. They are far less practised at forecasting quantities, because for a decade the quantity was headcount and headcount was already known.
The unit of price is no longer the seat
Salesforce moved Agentforce to Flex Credits in May 2025. Under the published model each credit costs $0.10, one action costs one credit, and the previous model charged $2 per conversation regardless of complexity. Bill Patterson, its EVP of CRM Applications, framed the change as pricing "on the value they deliver, not unrealized potential".
Look at the arithmetic rather than the framing. Salesforce's own example puts a simple inquiry at 3 to 6 actions, so $0.30 to $0.60 against $2 before. That is cheaper for simple work and more expensive for complex work, and it moves the risk of complexity from the vendor to you.
My view is that this is the correct pricing model and the harder one to buy. A per-seat contract is a fixed cost you can budget. A per-action contract is a variable cost you have to model, and most buyers have no history to model it from.
So ask for three volume scenarios in writing before signing: expected, double, and five times. A vendor confident in its unit economics will produce them. A vendor that resists is telling you something useful.
The discipline already exists elsewhere in the organisation. The FinOps Foundation's sixth annual survey, published in February 2026, covered 1,192 respondents representing more than $83 billion in annual cloud spend, and found 98% now manage AI spend against 31% two years earlier. The same survey found 90% now manage SaaS and 64% manage licensing.
That is the tell. The function built to forecast variable cloud spend has been handed the software portfolio, because software started behaving like cloud.
Data rights moved from the footnote to the redline
Security review used to mean a questionnaire and a certificate. AI purchases added a second question that the questionnaire was never designed to answer: what is the vendor allowed to do with what you put in?
Three clauses worth the fight
First, training use. Ask for a contractual commitment that customer data and outputs are not used to train or improve shared models, not just a settings toggle.
Second, output ownership. Assign ownership of outputs to the customer explicitly. Silence here is common, and silence favours whoever drafted the agreement.
Third, indemnity scope. Some AI agreements carve generated content out of IP indemnification entirely. If output infringes, the carve-out puts that risk on you, which is a strange place for it to sit given who chose the training data.
Buyers looking for drafting language have a free template. The US Office of Management and Budget issued memo M-25-22 on AI acquisition in April 2025. It directs agencies to include contractual terms setting out the ownership and IP rights of the parties. It tells them to review data ownership processes when government data is used to train AI. And it requires performance validation and pre-award testing for high-impact systems.
Strip out the government-specific parts and what remains is a serviceable commercial checklist.
The regulatory calendar adds a second pressure. Under Article 113, the remainder of the EU AI Act starts to apply on 2 August 2026, with Article 6(1) following on 2 August 2027. Certification is filling the gap in the meantime: Anthropic announced accredited ISO/IEC 42001 certification in January 2025, and the standard is appearing in vendor qualification questionnaires. Treat a certificate as evidence of process, and always ask for the scope statement, because scope is where these certifications differ.
How buyers are classifying tools as AI-replaceable
The most consequential change is quieter than any clause. Buyers are sorting their stack into categories of durability before the renewal calendar reaches them.
The four buckets
The sorting that shows up repeatedly in practice has four positions. Core platform, where the vendor holds the system of record. Essential extension, where the tool does something the platform genuinely cannot. Consolidation candidate, where a platform you already pay for now covers most of the need. And rebuild candidate, where the tool is a thin interface over data you already hold.
The fourth bucket is smaller than the discourse suggests and larger than vendors would like. The test I would apply is contractual rather than technical: what is the vendor accountable for when this fails at 3am, and what lands on your team if you rebuild it? Tools that survive that question are rarely the ones with the best features. They are the ones carrying real operational liability.
Writing for InformationWeek in July 2026, Aashis Luitel made the same point from the vendor side, arguing that the real cost of an internal rebuild is not development but support, security compliance, documentation and liability. That is a vendor-employed author making a self-interested argument, and it is also correct. The build-versus-buy arithmetic for agent-shaped tools is worked through in the piece on building or buying coding agents.
What Klarna actually did
The most-cited example of a buyer classifying SaaS as AI-replaceable is Klarna, and it is cited wrongly almost every time.
Sebastian Siemiatkowski's remarks about shutting down Salesforce and Workday were widely read as replacing SaaS with a language model. He corrected that directly, telling diginomica in March 2025: "So no, we did not replace SaaS with an LLM. Storing CRM data in an LLM would have its limitations." What Klarna did was consolidate, connect its data using Neo4j and other tools, and move some functions to alternative vendors including Deel.
The correction is more useful than the original story. The winning move was data consolidation, and the AI sat on top of it. That is a much harder programme to copy than cancelling a contract, and it is the same pattern behind the rationalisation work in portfolios running hundreds of applications.
A scorecard you can copy
Here is the artifact. Weights sum to 100, scoring is 1 to 5 per criterion, and the right-hand column is what ends the evaluation regardless of score.
| Criterion | Weight | What a passing answer looks like | Automatic fail |
|---|---|---|---|
| Cost at volume | 20 | Published unit rate plus a written forecast at expected, double and five times volume | Price quoted only per seat, no unit rate disclosed |
| Data rights and training use | 20 | Contractual no-training default, stated retention window, output ownership assigned to you | Training on customer data is a setting, not a term |
| Agent capability, evidenced | 15 | Named action inventory, permissions per action, documented failure behaviour | Agency demonstrated only in a scripted demo |
| Exit and portability | 15 | Full export of records, prompts, configuration and audit logs in a documented format | Export limited to a flat file of records |
| Model dependency | 10 | Named models, version pinning, notice period before a model changes | "We always use the best available model" |
| Governance evidence | 10 | Security attestation plus an AI management system, with the scope statement shared | Certification claimed, scope withheld |
| Integration depth | 5 | Works with your identity, logging and data layer without a bespoke build | Requires a parallel copy of your data |
| Support commitments | 5 | Contractual response times, named escalation, published deprecation policy | Commitments live only in the sales deck |
These weights are a starting position, not research. Re-cut them for your own risk profile before use.
The right-hand column matters more than the weights. My honest view is that the scorecard is worth less than its two hardest disqualifiers, which are the export test and the training-use test. Both are binary, both are cheap to check, and both predict how much freedom you will have in three years. Interoperability standards may eventually make the export question easier to answer, a shift covered in the analysis of MCP as an agent interoperability standard.
If you are on the other side of the table
Everything above is also a specification for what sellers now have to produce. The vendors clearing these reviews are not the ones with the best demo. They are the ones who arrive with the evidence already assembled.
| What the buyer asks | Answer that loses | Answer that survives |
|---|---|---|
| What does this cost at scale? | "It depends on your usage" | A rate card, plus three modelled volume scenarios you supply unprompted |
| What trains on our data? | "You can turn that off" | A contract clause, a retention window in days, and a deletion process |
| What does the agent actually do? | A recorded demo | An action list with permissions, error rates and an escalation path |
| What happens if we leave? | "We can arrange an export" | Documented export format, tested, with the audit trail included |
If I were selling into this process, I would publish the rate card. It costs a small amount of pricing flexibility. It removes the single largest source of stalled deals: a buyer who cannot forecast the bill cannot get it approved. Platform vendors face a sharper version of the same test, examined in the piece on whether Salesforce is a platform or roadkill in the agent era.
Where this argument is weakest
Three genuine problems with everything above, stated plainly.
The evidence base is thinner than the confidence
The strongest survey figure here comes from a venture capital firm's market report, reached through secondary coverage rather than a published methodology appendix. 141 CIOs is a real sample and a small one. The FinOps figures are better documented but describe practitioners in cost management roles, who are not a representative sample of buyers.
Nobody has published a large, methodologically transparent study of how enterprise software evaluation criteria changed between 2023 and 2026. Every claim in that direction, including mine, is an inference from adjacent data.
The number you should not repeat
Search for the state of AI procurement and you will meet a figure claiming that 88% of AI agent pilots never reach production. It appears in dozens of posts, attributed variously to IDC, Forrester, Anaconda, McKinsey and Gartner, sometimes several at once.
I could not trace it to a primary report during this research, and the attributions contradict each other. Six restatements of a figure are not six sources. If a vendor quotes it to you, ask which study, what sample and what date, and watch what happens.
Procurement rigour can become theatre
The final weakness is the scorecard itself. A weighted matrix produces a number, and a number produces confidence, and confidence is exactly what a buyer should not have when evaluating a category this young.
There is also a cost. A six-month evaluation of a tool with a 12-month useful life is a bad trade. The counter-argument to all of this is simple. Fast, reversible, small purchases beat careful, slow, large ones while the technology moves this quickly. That case is strong. I would apply the full scorecard only where switching costs are high, and buy small and cancel early everywhere else.
Frequently asked questions
How is software procurement changing in 2026?
Buying decisions now start with three questions that used to come last. What actions does the software take without a human? What does it cost per unit of work at production volume? And what rights do you keep over your data and outputs? The driver is budget, not fashion. Redpoint's March 2026 survey of 141 CIOs found 45% of AI spend is reallocated from existing software budgets rather than added to them.
What should be in an AI vendor evaluation scorecard?
Weight cost at volume and data rights highest, at roughly 20 points each. Then evidenced agent capability and exit portability at 15 each, model dependency and governance evidence at 10 each, and integration depth and support commitments at 5 each. Alongside the weights, define disqualifiers that end the evaluation regardless of score. The two most useful are refusal to publish a unit rate and training on customer data by default.
How do you forecast consumption-based AI pricing?
Get the unit rate in writing, then model three volumes: expected, double and five times. Pilot invoices understate production cost because pilots avoid retries, long context, failed tool calls and heavy users. Ask the vendor to supply the scenarios and check their assumptions against your own transaction counts. Salesforce, for example, publishes a rate of $0.10 per Agentforce action, which makes that arithmetic possible before signature.
What contract clauses matter most when buying AI software?
Three. A contractual commitment that your data and outputs are not used to train shared models, rather than a settings toggle. Explicit assignment of output ownership to you, because silence favours the drafter. And an IP indemnity without a carve-out for AI-generated content, since some agreements place that risk entirely on the customer. Add a documented export right covering data, prompts, configuration and audit logs.
Does the EU AI Act affect software buyers or only vendors?
Both. The Act places obligations on deployers as well as providers, so an organisation using an AI system inside its own operations carries duties of its own. Under Article 113 the remainder of the Act starts to apply on 2 August 2026, with Article 6(1) following on 2 August 2027. For buyers the practical effect is that your vendor evidence pack becomes part of your own compliance file.
Should we build internal tools instead of renewing SaaS?
Sometimes, and less often than the discussion implies. The build cost has fallen sharply. The support, security, documentation and liability costs have not. The honest test is what the vendor is contractually accountable for when the tool fails, and what lands on your team if you own it instead. Klarna's own CEO publicly corrected the idea that his company had replaced SaaS with a language model.
Where to start
Take the next renewal on your calendar and treat it as a first purchase rather than a continuation.
Before the call, write down the unit of work that vendor performs for you and how many units you consumed last quarter. If you cannot produce that number from your own systems, you are not ready to negotiate a consumption contract, and finding that out costs an afternoon rather than a contract term.
Then send two questions by email, in writing, before any meeting: what is your unit rate at our forecast volume, and what exactly can we export if we leave. The answers, and the time taken to produce them, will tell you more than the evaluation ever will.
Related analysis
The pricing side of this shift is covered in seat compression and what replaces per-seat pricing, and the market side in vertical AI eating horizontal SaaS.
References
- Redpoint Ventures, 2026 Market Update, March 2026. Survey of 141 CIOs. Used for the 45% budget reallocation, 54% consolidation and category replacement figures, as reported in SaaStr's summary of the survey.
- FinOps Foundation and the Linux Foundation, State of FinOps 2026 survey release, 19 February 2026. 1,192 respondents, more than $83 billion in annual cloud spend. Used for the 98%, 90% and 64% figures.
- EU Artificial Intelligence Act, implementation timeline and Article 113. Used for the 2 August 2026 and 2 August 2027 dates.
- Salesforce, Flexible pricing is a full circle moment for Salesforce, 23 May 2025. Used for the $0.10 per action rate, the prior $2 per conversation model and the Bill Patterson quote.
- Office of Management and Budget, M-25-22, Driving Efficient Acquisition of Artificial Intelligence in Government, April 2025. Used for the data ownership, IP rights and pre-award testing requirements.
- diginomica, Klarna: no, we didn't replace SaaS with an LLM, 7 March 2025. Used for the Siemiatkowski quote and the Neo4j consolidation detail.
- Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation prediction and the agent washing definition.
- Anthropic, Anthropic achieves ISO 42001 certification, 13 January 2025. Used as the example of AI management system certification entering vendor qualification.
Weakest part of this source base: the headline CIO figures come from a venture firm's market report, reached through secondary coverage rather than a published methodology appendix. The Gartner page returned an access error during research. Its figures are taken from Gartner's own published summary text rather than a page read end to end. Both are flagged where used.
Related reading