From Sanskriti Khandelwal | Product & Market Analysis

The AI Agent Procurement Checklist: Five Sections Most RFPs Leave Out

On this page

Gartner expects over 40% of agentic AI projects to be cancelled before the end of 2027, and it names the causes as cost, unclear value and weak risk controls. None of those three are feature problems. Your AI procurement checklist is measuring the wrong thing if it opens with a capability matrix. Capability is the one part of an agent purchase a vendor can demonstrate in an afternoon.

Key takeaways

  • Gartner counts roughly 130 genuine agentic AI vendors against thousands claiming the label. That ratio makes vendor verification the first job of an RFP, not the last.
  • Model retirement notice ranges from about 2 weeks to 6 months across the major providers. Anthropic commits to at least 60 days, OpenAI to at least 6 months for generally available models, and preview models can go with 14 days.
  • Consumption pricing moves the risk to you, and the billable unit is rarely defined in the RFP. Salesforce prices an Agentforce action at 20 Flex Credits, or $0.10, with a token allowance before it counts as two actions.
  • Ask for eval access, not eval results. A vendor scorecard measures the vendor's task. Only your own harness, run on your own data, measures yours.
40%+Of agentic AI projects Gartner expects to be cancelled by end 2027. Source: Gartner, June 2025.
130Agentic AI vendors Gartner judges to be real, out of thousands claiming the label. Source: Gartner, June 2025.
60 daysAnthropic's minimum retirement notice for a public model. Source: Anthropic documentation, 2026.

What belongs in an AI agent procurement checklist? Five sections: proof of capability on your own data, a map of where your data travels, a defined billable unit with a repricing history, contractual eval access, and exit terms covering model retirement. Features come last, because every shortlisted vendor already has them.

Why a feature-led RFP fails on agents

A traditional software RFP works because traditional software is deterministic. The same input produces the same output next quarter. You can therefore buy on a feature list and manage the rest through a service level agreement.

An agent breaks that assumption in three places. Its behaviour changes when the underlying model changes. Its cost changes with the work it is given. And its failure mode is a confident wrong answer rather than an error code.

So the questions that decide whether a deployment survives are not on the feature list. They sit in the contract, the billing schedule and the model lifecycle policy, which is exactly where a procurement team is least likely to look.

Gartner's own framing of the vendor market is blunt. It describes agent washing as the rebranding of assistants, robotic process automation and chatbots without substantial agentic capability. It estimates that only about 130 of the thousands of vendors using the label are real.

Treat that as an instruction about sequencing. If most shortlisted vendors are selling an old product under a new name, the first section of your evaluation has to be verification. Everything after it is wasted effort on a candidate that should never have made the list.

The cheapest verification is a single question. Ask the vendor to describe one decision the system makes without a human in the path, and what happens when that decision is wrong. Rebadged automation cannot answer the second half.

That leaves the feature matrix, which is the part every vendor is best at. I would cut it to a single page and move the freed time into the five sections below. A feature matrix mostly measures which vendor employs the better bid writer. The rest of this checklist measures what the thing costs you in month 14.

How much warning you get before a model disappears Minimum published notice before retirement, in days, per each provider's own documentation OpenAI, GA models180 OpenAI, variants90 Anthropic, public60 Foundry, GA60 Foundry, preview30 OpenAI, preview14 A pilot built on a preview model has a two week floor on its own continuity.
The bottom row is the one to check first. Preview models carry the shortest notice and are the ones a proof of concept is most likely to be built on.

Section 1: capability, proved on your data

Capability is worth 30 points on the sheet below, and none of them come from a demo. They come from a run on your material, scored by your people, against a task you were doing anyway.

Ask for a run on your data, not a demo

Give every shortlisted vendor the same 50 real cases from your own backlog, with the answers you actually produced. Ask them to run their agent and return the outputs plus the traces. Score the outputs blind, without knowing which vendor produced which.

This costs a week and it removes most of the guesswork. It also filters hard. A vendor who cannot ingest 50 of your documents inside a fortnight will not integrate with your systems in a quarter, and that is useful information on its own.

Fifty is a deliberate number. It is small enough that a vendor will agree to it and large enough that a 10 percentage point difference in accuracy is visible. Where a proof of concept has to become a production system, the wider set of tests appears in the readiness checks that separate a working demo from a deployable agent.

Ask what the agent does when it is wrong

Every agent is wrong sometimes. The question that matters is what the system does next, and whether you can see it happen.

Require three things in writing. A confidence or abstention path, so the agent can decline rather than guess. A trace for every action, retained long enough for an audit. And a defined escalation route to a human, with the trigger conditions named.

The most common failure I see is not a bad model. It is a good model wired so that nobody notices it drifted, which is the pattern catalogued in the breakdown of how agent pilots actually die.

Section 2: data handling, past the security questionnaire

Your standard security questionnaire was written for software that stores data. An agent sends data somewhere, and often to a party you have not contracted with.

Where do prompts, documents and logs actually go

Ask for a single diagram naming every destination for five artefacts: the prompt, the retrieved document, the embedding, the model output and the log. For each destination, ask which legal entity holds it, in which country, for how long, and whether it is used for training.

If the vendor cannot produce that diagram in a week, stop the evaluation. This is not a hard question for a company that has built the system it is selling. It is an impossible question for a company reselling somebody else's API behind a wrapper.

Pay particular attention to the sub-processor list, because that is where the model provider usually appears. The terms that apply once your data reaches that layer are compared in the review of what each foundation model provider commits to on training and retention.

The clause that can turn you into the provider

This is the part European buyers keep missing. Under Article 25 of the EU AI Act, a deployer or distributor can become the provider of a high-risk AI system. It happens in three ways: putting your own name or trademark on the system, substantially modifying it, or changing the purpose of a general purpose system so that it becomes high risk.

Becoming the provider means inheriting the provider obligations in Chapter III. That is a very different compliance position from the one your procurement team thinks it is signing up for.

Article 25 also gives you the remedy. Where a third party supplies a component, the original provider must specify by written agreement the information, capabilities and technical access needed to comply. Ask for that agreement by name during procurement. Asking for it after an incident is a negotiation you will lose.

High-risk obligations under the Act, including the Article 13 instructions for use and the Article 26 deployer duties, apply from 2 August 2026. The transparency-specific requirements are unpacked in the practical checklist for meeting the AI Act's disclosure rules.

Public buyers already have a template for this. The European Commission published updated model contractual clauses for AI procurement on 5 March 2025. There is a full version for high-risk systems, a light version for everything else, and a commentary. They are drafted for public organisations, and nothing stops a private buyer from lifting the clause language. It is better drafted than most vendor paper.

Section 3: cost predictability and the unit you are billed in

Consumption pricing is not inherently worse than a seat licence. It is inherently less predictable, and the RFP is where you decide who carries that uncertainty.

Salesforce is a useful worked example because its pricing is public. Under Flex Credits, announced on 15 May 2025, each Agentforce action consumes 20 credits, or $0.10, with credits sold in packs of 100,000 for $500. The previous model charged $2 per conversation.

Now read the definition rather than the price. An action covers a set token allowance, and work that exceeds it bills as more than one action. So your cost per task depends on how verbose your documents are, which is a property of your business and not of the vendor's price list. The arithmetic of that shift is worked through in the analysis of what a conversation actually costs under credit pricing.

Three questions settle this section. What is one unit, precisely. What causes one task to consume more than one unit. And what was your measured units-per-task on the 50-case run from section 1.

Then ask for the repricing history, because credit systems are repriced. The credit is an abstraction between you and the cost, and abstractions get adjusted. Ask every vendor for a written record of changes to the credit definition, the conversion rate and the included allowances over the past 24 months. A vendor with a clean history will send it. A vendor who cannot produce it is telling you something.

Then get two protections into the contract: a cap on rate changes during the term, and the right to convert to a fixed fee at a formula agreed now. The wider set of commercial terms worth fighting for is set out in the finance-side review of AI contract clauses.

The 18 month clock that starts before you sign Generally available model lifecycle on Microsoft Foundry, per Microsoft's published policy 0 Launch Retirement date set 12 months Deprecated Closed to new customers 15 Replacement named 90 to 120 days out 16 60 day notice 18 Retired 410 Gone Models from several partners run a 12 month lifecycle instead of 18. Retirement dates are not extendable.
The replacement model is not named until roughly three months before retirement. Any migration plan that assumes a known target earlier than that is planning against a blank.

Section 4: eval access, the clause nobody asks for

This is the section I would add if I could only add one. It is also the one most likely to be refused, which tells you how much it is worth.

Vendor benchmark numbers describe a task the vendor chose. Your task is different, your data is messier, and the gap between the two is where deployments fail quietly.

What you want in the contract is the right to run your own evaluation suite against the production system. Specify the frequency, and specify that the results are yours to keep and to use internally. Ask for a sandbox with representative latency, a stable test endpoint, and no rate limiting that makes a 500-case run impractical. Building the suite itself is a separate exercise, laid out in the guide to standing up an evaluation suite that reflects your own work.

Vendors resist this for a reason that is partly legitimate. Published third-party evals of a fast-moving system age badly and can be unfair. Concede the publication point and keep the access. Internal use is the part you actually need.

Then ask two mechanical questions. Can you pin the underlying model version, and how many days of notice do you get before a change to the model, the system prompt or the tool set. Thirty days is a reasonable floor for a production system. Anything under a fortnight means you will discover a behaviour change from a customer complaint rather than from a test.

Section 5: exit terms, including the exit you do not control

Procurement teams handle termination well. They handle deprecation badly, because deprecation is not a commercial event and does not appear in the standard template.

What the model providers themselves commit to

The published lifecycle policies are the floor under any agent product built on them, and they vary by an order of magnitude.

Published minimum notice before model retirement, August 2026
Provider and tierMinimum noticeWhat the policy says
OpenAI, generally available modelsAt least 6 monthsLongest published commitment of the three.
OpenAI, specialised variantsAt least 3 monthsCovers chat, Codex and deep research variants.
OpenAI, preview modelsAbout 2 weeksNamed in the policy as an example, not a floor.
Anthropic, publicly released modelsAt least 60 daysObserved 2026 windows ran 60 to 62 days.
Microsoft Foundry, GA modelsAt least 60 days active noticeRetirement date is set at launch, 18 months out.
Microsoft Foundry, preview modelsAt least 30 daysDeployments are force upgraded or terminated.

Sources are each provider's own published documentation, read on 29 August 2026. Partner-operated platforms such as Amazon Bedrock and Google Cloud set their own schedules, so the same model can carry different dates depending on where you run it. Policies change; re-read them at renewal.

The Anthropic record shows the policy working as written rather than generously. Three 2026 retirements ran 60, 61 and 62 days from notice to shutdown. Requests to a retired model fail. On Microsoft Foundry they return 410 Gone, and Microsoft states plainly that retirement dates are not extendable.

Two consequences for the contract. Your agent vendor should owe you notice at least as long as its own upstream provider owes it, and the cost of a forced migration should not land on you as a change request.

Export, and what you actually get back

Ask what leaves with you on termination and in what format. The list should include conversation and action logs, any fine-tuned artefacts you paid for, the evaluation results, the retrieval index or its source documents, and the configuration.

Then ask the harder question. How long does the vendor keep your data after termination, in every system including backups and any sub-processor. Get a number and a deletion certificate, not a policy reference. Where the failure has already happened and money is at stake, the remedies available are examined in the analysis of liability caps in agent contracts.

The scoring sheet, and how to weight it

Weights are where a checklist becomes a decision. Here is the allocation I would use for a first agent purchase in a regulated function.

A 100 point scoring sheet with features left off Proposed weighting for a first agent purchase in a regulated function 100 points Capability proved on your data30 Data handling and residency20 Cost predictability20 Eval access and version pinning15 Exit and deprecation terms15 Feature coverage0
This weighting is an editorial proposal, not measured data. The point of the zero on the last row is that features are a gate, not a score: a vendor either clears the threshold or leaves the shortlist.
The five sections, the one question, and the answer that should worry you
SectionThe question that does the workThe answer that should stop the process
Capability (30)Run our 50 cases and return outputs plus tracesA refusal, or a delay past three weeks
Data handling (20)Name every destination for prompts, documents, embeddings, outputs and logsNo diagram, or a sub-processor list that is not current
Cost predictability (20)Define one billable unit and show 24 months of repricing historyThe unit is defined by the vendor's own measurement with no audit right
Eval access (15)Can we run our own suite against production, quarterlyResults available only through the vendor's dashboard
Exit terms (15)What notice do we get on model change, and what leaves with usNotice shorter than the vendor's own upstream provider gives it
Rewriting three standard RFP questions
The usual questionThe version that produces evidence
Is your platform secure and compliant?Which legal entity processes our prompts, in which country, and for how long is each artefact retained?
What accuracy does your agent achieve?What did it achieve on the 50 cases we sent, scored blind by us, and what did it cost per case?
What is your pricing?What is one billable unit, what makes a task consume two, and how has that definition changed since 2024?

The left column is not worthless. It is a fast filter a non-specialist buyer can run, and it surfaces certifications that a specific question would skip. Use it as a screen, then replace it with the right column for the shortlist.

Where this checklist is weakest

A checklist that does not state its own limits is marketing. Three real ones.

The failure statistics circulating do not trace

You will see claims that 78%, 88% or 89% of agent pilots never reach production. I went looking for the underlying studies and could not open a primary source for any of them. Most trace to blog posts citing other blog posts.

The Gartner prediction used at the top of this post is a genuine published forecast from a named analyst, and it is still a forecast rather than a measurement. Treat all of these numbers as directional. If your board wants a failure rate, measure your own.

A checklist can be answered rather than satisfied

Any published set of procurement questions becomes a bid-response template within a quarter. Vendors will have polished answers to all five sections before long, and polished answers are not evidence.

That is why the capability section is built on a run rather than a response, and why the cost section asks for history rather than a policy. Where a question can be answered with prose, it will be. Design around that.

It is calibrated for a first purchase, not a portfolio

The weighting above suits a buyer making an early agent purchase in a function that carries regulatory or financial risk. A team on its fourth deployment should shift weight away from data handling, which is by then a solved architectural question, and toward integration depth and cost.

The exit section is also the least popular and the easiest to trade away. If you have one concession to spend, spend it here rather than on price. The price is renegotiable at renewal and the migration is not.

Frequently asked questions

What should be in an AI agent procurement checklist?

Five sections. Capability proved on your own data rather than a demo. A written map of where prompts, documents, embeddings, outputs and logs travel. A defined billable unit with a repricing history. Contractual access to run your own evaluations against production. And exit terms that cover model deprecation as well as termination. Features belong in a pass or fail gate rather than a scored section, because every shortlisted vendor will clear them.

What questions should you ask an AI agent vendor?

Ask the vendor to run 50 of your real cases and return the outputs with traces. Ask which legal entity processes your data, in which country, and for how long. Ask what one billable unit is and what makes a single task consume two. Ask whether you can pin a model version. Ask how many days of notice you get before the model, prompt or tool set changes.

How much notice do AI vendors give before retiring a model?

It varies widely. OpenAI publishes at least 6 months for generally available models, at least 3 months for specialised variants, and about 2 weeks for preview models. Anthropic commits to at least 60 days for publicly released models, and its 2026 retirements ran 60 to 62 days. Microsoft Foundry gives at least 60 days for generally available models and 30 days for preview models.

How do you evaluate AI vendors for data security?

Go past the questionnaire and ask for a diagram naming every destination for five artefacts: the prompt, the retrieved document, the embedding, the model output and the log. For each destination, establish the legal entity, the country, the retention period and whether the data is used for training. Check the sub-processor list, because the foundation model provider usually appears there under different terms than the vendor's own.

What are the key contract terms for buying an AI agent?

A cap on consumption rate changes during the term. The right to run your own evaluation suite against production. A model version pinning option. A minimum notice period for model changes that matches the vendor's own upstream commitment. And a defined export package at termination with a deletion certificate. European buyers should also secure the Article 25 written agreement covering information and technical access.

Is an AI agent RFP different from a normal software RFP?

Yes, in three ways. Agent behaviour changes when the underlying model changes, so version and notice terms matter more than feature parity. Cost varies with the work performed, so the billable unit needs defining rather than pricing. And failure looks like a confident wrong answer rather than an error, so tracing, abstention and escalation paths have to be contracted rather than assumed.

Where to start this week

Pick one live evaluation and do two things to it.

First, assemble the 50 cases. Pull real work from the last quarter, with the answers your team produced, and send the same set to every shortlisted vendor. That single artefact does more to separate vendors than the rest of the RFP combined, and it takes an afternoon to build.

Second, open the model lifecycle page of whichever foundation model sits underneath each vendor's product. Write the retirement dates into your own risk register. You now know something about the continuity of that product that the vendor's own sales team probably does not.

Related on buying and pricing

The commercial half of this checklist is covered in more depth in the finance-side review of AI contract clauses and in the piece on how AI-assisted buyers now build a shortlist.

References

  1. Gartner, Gartner predicts over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation forecast, the agent washing definition and the estimate of about 130 genuine vendors.
  2. Anthropic, Model deprecations, read 29 August 2026. Used for the 60 day minimum notice and the 2026 retirement dates.
  3. OpenAI, Deprecations, read 29 August 2026. Used for the 6 month, 3 month and 2 week notice tiers.
  4. Microsoft, Foundry Models lifecycle and support policy, updated 24 July 2026. Used for the 18 month lifecycle, the 60 and 30 day notice periods, the 90 to 120 day replacement window and the 410 Gone behaviour.
  5. EU Artificial Intelligence Act, Article 25, Responsibilities along the AI value chain. Used for the conditions under which a deployer becomes a provider and for the paragraph 4 written agreement.
  6. European Commission Public Buyers Community, Updated EU AI model contractual clauses now available, 5 March 2025. Used for the existence and structure of the MCC-AI templates.
  7. Salesforce via Business Wire, Salesforce introduces new flexible Agentforce pricing, 15 May 2025. Used for the Flex Credits figures.

Weakest thing about this source base: the scoring weights in this post are an editorial judgement, not a measured result, and no public dataset ranks agent deployment failures by cause. The provider lifecycle policies are primary and current as of 29 August 2026, but every one of them can be revised by the provider without notice to buyers.

SK
Sanskriti Khandelwal
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading