From Sanskriti Khandelwal | Product & Market Analysis

When to Kill an AI Project: Six Criteria to Write Before You Start

On this page

42% of companies scrapped most of their AI initiatives in 2025, up from 17% a year earlier. Almost none of those stops were planned. Kill criteria are the thresholds you write before an AI project starts, so that stopping becomes a reading rather than an argument. This post gives six of them, with numbers attached.

Key takeaways

  • Abandonment is the base rate, not the exception. The share of companies scrapping most AI initiatives rose to 42% in 2025 from 17% the year before, across more than 1,000 respondents in North America and Europe.
  • The stop decision arrives late because nobody wrote it down early. Escalation research in software projects has shown for decades that money already spent, rather than progress made, drives the choice to continue.
  • Six thresholds cover most of the failure surface. Baseline, quality plateau, unit cost, intervention rate, adoption and named ownership. Each one is a number you can read on a Monday morning.
  • A kill and a redirect are the same family of decision. Treating every tripped threshold as a termination is what teaches teams to hide the reading rather than report it.

The short answer

Kill an AI project when a threshold you wrote in advance is breached and the team cannot explain the breach with new evidence. The working set is six. No recorded baseline, a quality score that has stopped moving, and unit cost above the method you already use. Then an intervention rate that will not fall, adoption under 20% at week 8, and no named owner.

42%Companies that scrapped most of their AI initiatives in 2025, up from 17% in 2024. Source: S&P Global Market Intelligence, reported by CIO Dive, March 2025.
40%+Share of agentic AI projects Gartner predicts will be cancelled by the end of 2027. Source: Gartner, June 2025.
71%CIOs who say they have until mid-2026 to prove AI value or face budget and job consequences. Source: Harris Poll for Dataiku, February 2026.

What a kill criterion is, and what it is not

A kill criterion is a condition, agreed in advance, that ends a project when it is met. It is not a warning sign. It is not a risk register entry. It is the sentence that decides.

Most AI governance packs contain success criteria and nothing else. That asymmetry is the defect. A success bar tells a team when to celebrate. It says nothing about when to stop, and stopping is the decision that costs real money to get wrong.

The point of writing the criterion early is not precision. It is that you write it while you are still indifferent to the answer.

A criterion is a number, a date and an owner

The number is a measurement you can read off a system you control. The date is when someone reads it. The owner is the person who has to say the result out loud in a room with the sponsor in it.

The missing part is almost always the owner. A threshold nobody owns is read by nobody. A project with no scheduled reading never fails a test it never sits.

Write all three parts or write none. Two out of three produces a document that makes the governance slide look finished and changes no behaviour at all.

Kill and redirect are not the same decision

Mark Keil and Daniel Robey studied troubled software projects in 1999, in the Journal of Management Information Systems. They defined de-escalation as the reversal of commitment to a failing course of action, either through project termination or through redirection. Both count. Both are wins.

That distinction matters operationally. If every tripped threshold means termination, the team will fight the reading rather than report it. Make redirection the standard response and the readings get honest.

My view is that a governance process which cannot produce a redirect is not governance. It is a firing squad with a calendar.

Abandonment is already running at 42%, and it arrives late

The base rate is not a secret. S&P Global Market Intelligence found the share of companies scrapping most of their AI initiatives rose to 42% in 2025, from 17% a year earlier, across more than 1,000 respondents in North America and Europe. The same work put the average share of proof of concepts scrapped before production at 46%.

Gartner predicts that over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. Its earlier prediction on data readiness is blunter still: through 2026, organisations will abandon 60% of AI projects that are not supported by AI-ready data.

RAND interviewed 65 data scientists and engineers across government and industry and reported that more than 80% of AI projects fail, roughly twice the rate of other IT projects. Read that figure carefully. RAND presents it as an estimate it is repeating rather than one it measured, and that hedge belongs in every citation of it.

Abandonment more than doubled in a single year Self-reported by more than 1,000 companies in North America and Europe SCRAPPED MOST AI INITIATIVES 17% 2024 42% 2025 PROOF OF CONCEPTS SCRAPPED 46% 2025 average Source: S&P Global Market Intelligence, reported by CIO Dive, March 2025.
Note what these numbers describe. They count projects that were abandoned, not projects that were stopped on a scheduled date against a written test.

That distinction is the whole point of this piece. None of the figures above record a stop decision. They record collapses, written up afterwards. A project that is abandoned ran out of sponsor patience. A project that is killed failed a test.

The difference is money and credibility. A project stopped at a gate returns its unspent budget and its team. An abandoned project returns neither, and it poisons the next approval. Where measurable return has actually shown up is a separate question, covered in the piece on who is making money from generative AI.

Why sunk cost holds AI projects longer than other IT projects

Late stops are not caused by ignorance. They are caused by a bias with 40 years of literature behind it, plus two pressures specific to this cycle.

The escalation research is old and unambiguous

Keil's experimental work on software project escalation found that the level of sunk cost influenced a decision maker's willingness to continue a troubled project. It also found that people justify continuing on the basis of money already spent more readily than on how much of the work is finished. Later replication found the effect held across cultures.

Estimates inside that literature put the share of organisational IS projects showing some degree of irrational persistence at 30% to 40%. That was measured long before anyone had an AI budget to defend.

AI adds two pressures the old studies did not model

The first is the demo. An AI project produces an impressive artefact within days. A working demo is evidence of capability. It is not evidence of value inside a workflow, and the two get conflated in almost every steering committee.

The second is the narrative. Harris Poll surveyed 600 CIOs at companies above $500 million in revenue, fielding the work for Dataiku between December 2025 and January 2026. 95% already brief their board on AI performance and 46% do so at least monthly. 74% said they regret at least one major AI vendor or platform decision made in the previous 18 months.

Put those together with the 71% who believe they have until mid-2026 to prove value. The project has been shown to the board, the CIO's pay is tied to AI outcomes, and the clock is short. That is an escalation machine with a governance label on it.

I would treat any project that has been demoed to the board twice without a recorded baseline as the highest risk item in a portfolio. Not the most likely to fail technically. The most likely to survive well past the point it should have stopped.

Six kill criteria, with thresholds you can actually set

Here are the six. The thresholds are defaults. Set them once, in writing, at approval, and then leave them alone until the reading date.

Six kill criteria and their default trip thresholds
CriterionTrips whenRead it from
1. No recorded baselineNo pre-deployment measurement of the target metric exists by gate 1.Your own logs, taken before the tool arrives.
2. Quality plateauHeld-out evaluation score moves less than 2 points across two consecutive build cycles.An eval set built before the first model call.
3. Unit costCost per completed task sits above the current method and has not fallen across two monthly readings.Vendor invoice plus your own usage logs.
4. Intervention rateHuman review or rework is still above half the level the business case assumed, at week 12.Queue and ticket data.
5. AdoptionWeekly active use inside the target group is under 20% at week 8, with no upward trend.Product telemetry you own.
6. Accountable ownerNo named person owns production incidents by gate 3.The org chart and the on-call rota.

These threshold values are editorial defaults chosen for being cheap to measure and hard to argue about. They are not measured optima, and no published dataset establishes the optimal values. Expect to move them after two cycles.

Two of the six get argued about, so here is the reasoning behind them.

Unit cost is measured against the method you already run, at the volume you actually run it. Not against a vendor benchmark, and not against a future price. Token prices keep falling while bills keep rising, a pattern set out in the analysis of cheap tokens and rising AI bills.

Adoption is the criterion teams resist hardest. A tool used by 15% of its target group at week 8 is not early. It is a workflow rejecting a tool. Which functions actually pay back, and over what period, is worked through in the payback-by-function breakdown.

Criterion 2 has a dependency worth naming. It cannot be read without a held-out evaluation set that existed before the first build. If you do not have one, criterion 1 has already tripped. Building that suite is its own job, covered in the guide to building an internal eval suite.

Where to set the numbers, and who holds the pen

Set every threshold on an instrument you own. Vendor dashboards are the wrong source. They report what the vendor chose to count, and the definitions change without notice or version history.

Prefer absolute numbers wherever the denominator can move. "Under 200 completed tasks a week" survives a team reorganisation. "Under 20% adoption" does not, because someone will quietly redefine the target group the week before the reading.

Give every threshold a date in the calendar before work starts. A reading that has to be requested does not happen. A reading already in three diaries happens even when the news is bad.

Whoever wrote the business case should not hold the pen

The de-escalation research found that a change of management, or a rotation of duties, reduces commitment in an escalation situation. That is not a comment on anyone's integrity. It is a comment on how humans hold positions they have argued for in public.

So split the roles. The sponsor argues the case and runs the work. A second person, with budget authority and no authorship of the original case, reads the thresholds and records the result.

This is also the honest answer to the centre of excellence question. A central team that owns the readings and not the projects is useful. One that owns both becomes the bottleneck described in the piece on measuring decision latency in an AI CoE.

The gate structure that makes a kill cheap

Kill criteria without gates produce arguments. Gates give the reading somewhere to happen and someone to happen in front of.

Robert Cooper's stage-gate system, built for physical product development, treats gates as go or kill investment points where weaker projects are culled at each successive stage. The mechanism transfers directly. What changes for AI is the spacing, because the cost curve is steeper and the first real evidence arrives faster.

Four gates, and the cheapest place to stop Each gate is a scheduled reading of the thresholds, not a status update Gate 1 Week 0 Gate 2 Week 4 Gate 3 Week 8 Gate 4 Week 26 Baseline exists? Criterion 1 Eval score moving? Criterion 2 Adoption and cost Criteria 3 and 5 Owner and incidents Criteria 4 and 6 Cost at risk grows from left to right. Gate 1 is the only free stop. Gate spacing is an editorial default. Adjust it to your build cycle, not to your quarter.
Gate 1 costs a meeting. Gate 4 costs a year of integration work. The criteria are the same at each one, and only the price of ignoring them changes.

Four gates are enough. Gate 1 asks one question: is there a recorded number this project is meant to move? Gate 2 reads the evaluation score. Gate 3 reads adoption and unit cost inside a real workflow. Gate 4 reads the incident path and the named owner before any renewal is signed.

Run a pre-mortem at gate 1. Gary Klein's version, published in Harvard Business Review in 2007, asks a team to assume the project has already failed and to write down independently why. It runs in under half an hour. Klein cited 1989 work by Deborah Mitchell, Jay Russo and Nancy Pennington finding that prospective hindsight raised the ability to identify reasons for a future outcome by 30%.

The output of that session is your first draft of the kill criteria. The failures a team predicts on day one are the ones worth setting thresholds against, and they are usually more specific than anything a template will give you.

What happens on the day a criterion trips

A tripped threshold starts a decision. It does not execute one. The question is which of three moves applies, and the answer comes from two readings, not one.

The three responses to a tripped threshold
ResponseApplies whenWhat it looks like in practice
Narrow the scopeValue evidence is real but confined to a subset of cases.Cut the use case list. Keep the cases that clear the bar, drop the rest.
RedirectThe cost curve is improving but the current problem is the wrong one.Same team, same stack, different workflow. The build is reused.
KillBoth readings are bad and no new evidence has arrived since the last gate.Stop the spend, write the reading down, harvest the assets.
Two readings decide it, not one Read value evidence and unit cost trajectory on the same date Continue Fund the next gate. Nothing else. Narrow the scope Keep only the cases that clear the bar. Redirect Same team and stack, new workflow. Kill Stop the spend. Harvest the assets. Value evidence present Value evidence absent Unit cost falling Unit cost flat or rising Illustrative framework, not measured data. The axes are the two readings, not scores.
Only one quadrant is a kill. That ratio is deliberate, because a process that produces nothing but terminations will stop receiving honest readings.

The clearest public example is a redirect, not a kill. Klarna built a customer service assistant that handled roughly two thirds of inquiries, then rebuilt the operating model around human agents for complex work. Chief executive Sebastian Siemiatkowski said cost had been "a too predominant evaluation factor" and that lower quality was the result, in comments reported in May 2025.

Read what actually happened there. The assistant stayed on routine chats. The scope narrowed and humans returned for disputes and complaints. That is criterion 4 tripping on quality and intervention, answered with a scope reduction rather than a shutdown.

Whichever move you make, harvest the assets before the team disperses. A killed AI project usually leaves behind four things worth keeping: the evaluation set, the recorded baseline, the cleaned data pipeline, and the integration work into the system of record. Those are the expensive parts. The model was never the expensive part.

Write the reading down in a file with the date on it. Your own stopped projects are the only AI case studies you can fully verify, and the difficulty of verifying anyone else's is the subject of the forensics piece on published AI ROI claims.

Where this argument is weakest

Three problems with everything above, stated plainly rather than buried in a footnote.

The thresholds here are defaults, not measured optima

We hold no first-party data on which threshold values maximise portfolio return, and no published dataset settles it either. The six criteria are chosen because each is cheap to measure and hard to argue with at a gate. The specific numbers are editorial judgement. Anyone presenting numbers like these as validated is overclaiming.

Stopping too early has a cost nobody bills you for

A troubled project can be an option on a capability that becomes valuable later. Killing at week 8 on a flat evaluation score can destroy that option, and the loss never appears in any report, because the counterfactual is invisible.

The bias in this piece runs toward stopping. That is a choice, not a finding. If your process produces only kills and never redirects, it is over-tuned, and you will lose work that was 6 weeks from functioning.

The evidence base is self-reported surveys

The 42% and the 40% both come from surveys. Respondents report abandonment against their own definitions, sample frames differ, and vendors commission a large share of this research, including the CIO study quoted above. RAND's figure is an estimate it repeats rather than one it measured.

Treat all of it as directional evidence that stopping is common and unplanned. Do not treat any of it as a measured failure rate for your own portfolio.

Frequently asked questions

When should you kill an AI project?

Kill it when a threshold you wrote before the work started has been breached and the team cannot explain the breach with new evidence. The useful thresholds are a missing baseline, a quality score that has stopped moving, and unit cost above the method you already use. Add an intervention rate that will not fall, adoption under 20% at week 8, and no named owner for production incidents.

What are kill criteria in project management?

Kill criteria are pre-agreed conditions that end a project once they are met. Each one carries three parts: a number, a date on which the number gets read, and a person who has to state the result. They differ from success criteria because they describe the exit rather than the target, and they are written before anyone has become invested in the answer.

How do you avoid the sunk cost fallacy on AI projects?

Separate the person who wrote the business case from the person who reads the threshold. Research on software project escalation found that decision makers justify continuing on money already spent rather than on progress made, and that rotating the responsible manager reduces that pull. Write the exit conditions at approval, read them on a fixed date, and record the reading whether or not it is comfortable.

How long should an AI pilot run before you stop it?

Set the clock in build cycles rather than months. A workable default is four gates: week 0 for the baseline, week 4 for the evaluation score, week 8 for adoption and unit cost, and renewal for ownership. A pilot that has passed 12 weeks with no named production owner is not a pilot any more. It is an unfunded product with a sponsor.

What percentage of AI projects get cancelled?

S&P Global Market Intelligence found that 42% of companies scrapped most of their AI initiatives in 2025, up from 17% a year earlier, across more than 1,000 respondents in North America and Europe. The same work put the average share of proof of concepts scrapped before production at 46%. Gartner separately predicts that over 40% of agentic AI projects will be cancelled by the end of 2027.

Should you kill an AI project or redirect it?

Redirect when the cost curve is improving but the problem is wrong, and kill when both readings are bad. Research on de-escalation treats termination and redirection as the same family of decision, being the reversal of commitment to a failing course. Klarna's customer service deployment is the public example of a redirect, since the assistant stayed on routine chats while humans returned for complex cases.

Where to start this week

Two things, both doable before the next steering meeting.

First, take your largest live AI project and write its six thresholds on one page. Put a date and a name against each one. If you cannot fill in the baseline row, you have your first finding, and you did not need a gate to get it.

Second, book the gate 2 reading in the calendar now, with the person who did not write the business case as its owner. Then leave the numbers alone until that date arrives. The value of a threshold comes entirely from having been set while you did not know the answer.

Related

The other half of this question is what a project costs while it stays alive. See the analysis of AI deployments with negative measured return.

References

  1. CIO Dive, AI project failure rates are on the rise, 14 March 2025. Used for the S&P Global Market Intelligence abandonment and proof of concept figures.
  2. Gartner, Over 40% of agentic AI projects will be canceled by end of 2027, 25 June 2025. Used for the cancellation prediction and its three stated causes.
  3. Gartner, Lack of AI-ready data puts AI projects at risk, 26 February 2025. Used for the 60% abandonment prediction through 2026.
  4. RAND Corporation, The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed, 2024. Used for the 65 interviews and the hedged 80% figure.
  5. Harris Poll for Dataiku, 71% of CIOs say they have until mid-2026 to prove AI value, 12 February 2026. Used for board reporting, regret and deadline figures.
  6. Keil and Robey, Turning Around Troubled Software Projects, Journal of Management Information Systems, Vol 15 No 4, 1999. Used for the de-escalation definition and the management rotation finding.
  7. Gary Klein, Performing a Project Premortem, Harvard Business Review, September 2007. Used for the pre-mortem protocol and the 30% prospective hindsight finding.
  8. CX Dive, Klarna changes its AI tune and again recruits humans for customer service, May 2025. Used for the Siemiatkowski quotation and the deployment scope.

The weakest thing about this source base: the two headline abandonment figures are self-reported survey results, with no common definition of what counts as scrapping a project. One of the CIO studies quoted was also commissioned by a software vendor. The threshold values in the six-criteria table are editorial defaults and carry no supporting dataset at all.

SK
Sanskriti Khandelwal
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading