From Sanskriti Khandelwal | Product & Market Analysis

AI-Ready Data Blocks 76% of Agent Deployments. Here Is the Repair Order

On this page

76% of enterprise data leaders name data quality and trust as a top-three barrier to putting AI agents into production. Only 8.4% say the data feeding their AI systems is trustworthy enough for production use. The gap between those two numbers is what AI-ready data work has to close, and closing it has a correct order that most programmes get backwards.

Key takeaways

  • 75.9% of data leaders call data quality and trust a top-three blocker to production agents. In the same survey of more than 540 practitioners across 66 countries, missing context and lineage came second at 63.5% and security third at 61.7%.
  • Only 8.4% trust their data enough for production AI, and 57.3% are running or piloting agents anyway. Most agent deployments are therefore running on data their own custodians will not vouch for.
  • Gartner expects 60% of AI projects unsupported by AI-ready data to be abandoned through 2026. 63% of organisations told Gartner they either do not have, or are unsure they have, the data management practices AI requires.
  • Fixing data horizontally is its own documented failure mode. Gartner separately predicts 80% of data and analytics governance initiatives will fail by 2027, which is why the sequence that works is scoped to one workflow rather than the estate.
75.9%Name data quality and trust a top-three barrier to production AI agents. Source: The Modern Data Company survey, August 2026.
8.4%Say the data feeding their AI systems is trustworthy enough for production. Source: The Modern Data Company, August 2026.
60%Of AI projects unsupported by AI-ready data will be abandoned through 2026, per Gartner, February 2025.

What "AI-ready data" actually means, and what it does not

AI-ready data is data qualified for a specific AI use case rather than certified in general. Gartner describes it as data aligned to that use case, governed at the asset level, supported by pipelines with quality gates, described by live metadata, and continuously quality-assured. It is not a grade you award a warehouse.

Melody Chien, Senior Research Director in Data Management at Gartner, put it plainly to TechTarget in March 2026: "AI-ready data means that data is ready to support certain AI cases." There is no universal label available.

Traditional data quality asked whether a field was accurate, complete and consistent. That question has one answer per field, and you can answer it once for the whole business. AI readiness asks something else entirely: whether a field carries enough context for a system to act on it without a person checking the result. That question has a different answer for every workflow you point at it.

A customer record good enough for a quarterly revenue report can be useless to a renewal agent that needs to know which contract governs which subsidiary. Both read the same table. Only one of them is about to send an email.

Unstructured data is most of the pile

Between 70% and 90% of enterprise data is unstructured, according to the Gartner analysts TechTarget interviewed. Only 32% of organisations with AI initiatives have any data readiness process at all, which leaves 68% doing it unsystematically or not doing it.

So the largest share of the material your agents read is also the share least covered by conventional data quality tooling, which was built for rows and columns. Profiling a table tells you nothing about a four-year-old policy document that contradicts the current one.

The blocker, measured

In August 2026 The Modern Data Company published its annual survey of enterprise data leaders and practitioners, drawing more than 540 responses across 66 countries.

Barriers preventing AI agents from reaching production.
BarrierShare naming it top threeWhat it looks like in practice
Data quality and trust75.9%Nobody will sign off that the numbers are right.
Missing context and lineage63.5%The agent cannot say where an answer came from.
Security concerns61.7%Access scope is broader than the task requires.
Skills gap25.9%No owner for the data an agent depends on.
Immature tooling19.5%Platform limits, cited least often of the five.

Respondents selected three barriers, so the column sums above 100%. The survey was run by a data infrastructure vendor among self-selected respondents. The practice column is interpretation, not survey wording.

The adoption side of the same survey explains the urgency. 57.3% are piloting or running agents in data and analytics workflows, split between 23.5% in production and 33.8% in pilots. Set that against 8.4% who consider the underlying data trustworthy enough for production, and you have the actual shape of the problem. Deployment did not wait for readiness. It never does.

What stops agents reaching production Share of 540+ enterprise data leaders naming each barrier in their top three, 2026. Data quality and trust75.9% Missing context, lineage63.5% Security concerns61.7% Skills gap25.9% Immature tooling19.5% The two lightest bars are the ones vendors sell against. The three darkest are internal work. Source: The Modern Data Company, 66 countries, August 2026. Respondents chose three.
Notice what is not in the top three. Model capability and tooling maturity rank last, which is a finding about where budget should go.

The abandonment figures people quote

Two numbers circulate together and should not. Gartner predicted in February 2025 that through 2026, organisations will abandon 60% of AI projects unsupported by AI-ready data. That rests on a Q3 2024 survey of 248 data management leaders, 63% of whom lacked or were unsure of their practices.

Separately, S&P Global Market Intelligence found the share of businesses scrapping most of their AI initiatives rose to 42%, from 17% a year earlier. That work covered more than 1,000 respondents in North America and Europe, and it found the average organisation scrapped 46% of proof-of-concepts before production.

These get quoted as though they corroborate each other. They do not. One is a forecast about a subset of projects defined by their data. The other is a survey result about abandonment from any cause, and the top obstacles S&P respondents named were cost and privacy rather than quality. Stacking them produces a bigger number and a worse argument. The breakdown of where AI deployments actually return nothing takes that measurement problem apart in more detail.

Why agents break on data that dashboards tolerated

Your data did not get worse when you bought an agent. What changed is that the error-absorbing layer, which was a person reading a chart, was removed from the middle of the workflow.

A dashboard shows a number to someone who knows the business and notices when it looks wrong. An agent takes the number and acts on it, then takes its own output as the input to the next step.

A field that is correct 99% of the time survives a report comfortably. Across a 20 step chain it becomes a different object: the chance that every step reads a clean value falls to roughly 82%. That is arithmetic, not a study, and it is why a tolerable defect rate at the field level turns into a 1-in-5 failure rate at the workflow level. The specific ways that plays out are catalogued in the analysis of how agent pilots actually fail.

Agents read what nobody curated

Your warehouse was curated. The wiki, the ticket history, the shared drive and four years of email that your retrieval index now covers were not.

Those sources carry the context an agent needs. They also carry every superseded policy, draft price list and abandoned naming convention your organisation has ever produced. Retrieval does not know which document is current. It knows which document is similar, and a rescinded 2023 discount policy is extremely similar to the discount policy.

The most common thing I would expect to find inside a stalled deployment is not a weak model. It is a retrieval corpus that nobody has ever deleted anything from. Access scope compounds it, which is the failure mode examined in the comparison of enterprise assistant permission models.

Step 1: scope the repair to one workflow, not the warehouse

Everything in this sequence depends on getting this step right. It is also the step most programmes skip on the way to a platform purchase.

Gartner predicts that 80% of data and analytics governance initiatives will fail by 2027, and the stated cause is that they are not tied to a prioritised business outcome. Saul Judah, the VP Analyst behind that prediction, framed it directly: a governance programme that does not enable prioritised business outcomes fails.

Practitioners working on enterprise data modernisation describe the alternative as refusing to boil the ocean. Pick two or three use cases with clear business value and let those drive the architecture.

The two questions that scope the work

Write the workflow down, then answer two questions about it.

First, which fields does the agent read, and which does it write? Second, what is the worst thing it can do with a wrong value in each one? The first question produces a list that is almost always shorter than the team expected, usually a few dozen fields rather than a schema. The second sorts that list into the order you fix it in.

Field scoping worksheet, illustrative example for a renewal agent.
FieldRead or writeWorst case if wrongVerdict
contract_end_dateReadAgent renews or lapses an account on the wrong day.Fix first
entitlement_tierReadAgent quotes a price the contract does not support.Fix first
account_ownerReadEscalation routes to someone who left the company.Fix first
last_ticket_sentimentReadTone of the outreach is slightly wrong.Accept as is
crm_activity_noteWriteInternal record is noisy, no external effect.Accept as is

This table is illustrative and shows the shape of the exercise, not measured data from any deployment. The point is the last column: most fields an agent touches do not need remediation, and finding out which is the cheapest hour in the project.

Everything downstream in this sequence is cheap once that list exists, and close to impossible before it.

Step 2: fix identity and freshness on the fields that workflow touches

The first repair is almost never accuracy. It is identity.

Suppose one customer exists as three rows across the CRM, the billing system and the support desk. Every number the agent computes about that customer is then wrong in a way no field-level quality check will catch. Each row is internally valid. Every field passes. The entity is still fictional.

Resolve entities first, enrich second. Teams reverse this constantly, buy an enrichment product, improve completeness on all three duplicate records, and then wonder why the cleaner fields changed nothing.

Freshness on the fields that get quoted

Decide a staleness budget for each field and enforce it in the pipeline. A contract end date read from a nightly sync is fine. A credit limit read from a nightly sync is not, if the agent can commit you to something between syncs.

Attach a timestamp to every field the agent is permitted to quote, and have the agent refuse to answer rather than answer from outside the budget. An agent that says it does not know is far cheaper than an agent that is confidently 14 hours out of date, and the second failure is the one that reaches a customer.

Step 3: give the agent context, not just columns

Gartner was unusually direct about this layer in May 2026. The press release ran under the headline that a lack of semantics causes inaccurate AI agents and wasted spending. It forecast that by 2027, organisations prioritising semantics in AI-ready data will raise agentic AI accuracy by up to 80% and cut costs by up to 60%. Rita Sallam, Distinguished VP Analyst, framed context with semantic coherence as a cost control and trust strategy rather than an optional extra.

Read the direction of that forecast and ignore the magnitude. "Up to 80%" is doing a great deal of work in a sentence about a period that has not happened yet, and no methodology accompanies it publicly. The claim that semantics matters is well supported. The claim that it is worth precisely 80% is not.

"Active customer" has three definitions inside most companies: one in finance, one in sales, one in the product analytics tool. A human reconciles them without noticing they differ. An agent picks whichever definition its retrieval happened to surface, and it picks a different one next Tuesday.

Write one definition for each term the workflow depends on. Put it where the agent reads, not in a governance wiki nobody indexed. Record which system is authoritative for each term, because that is the question every dispute about an agent's output eventually reduces to.

Make the answer traceable

65.1% of the same survey's respondents said their AI use requires explainability and traceability. 39% maintain audit trails or data source links, and only 10% maintain both. That gap is the practical meaning of lineage, and it is a bigger operational problem than it sounds.

If your agent cannot say which record produced a number, you cannot debug it when it is wrong, and you cannot defend it when a regulator or a customer asks. Lineage is not a compliance artifact here. It is the only debugging tool you will have at 3am.

Agents are already deployed. The data is not trusted. Share of 540+ enterprise data leaders, 2026. 57.3% 23.5% live Piloting or running agents 8.4% Trust the data for production The gap 49 points of unbacked deployment Source: The Modern Data Company, 66 countries, August 2026.
The blue bar is a decision already made. The red bar is the evidence base under it. Most readiness work is done after this picture exists, not before.

Step 4: gate the pipeline, then watch what moves

A quality gate that logs a warning is a quality gate that does nothing.

The pipeline feeding an agent should stop when a check fails, and the agent should degrade to refusing the task rather than proceeding on data that failed validation. This is unpopular internally because it converts silent wrongness into visible downtime, and visible downtime generates tickets. I think that trade is obviously correct: silent wrongness in an agent workflow is not avoided, it is simply discovered later, by a customer, in public. The tests worth running before this point are set out in the production readiness checklist for AI proofs of concept.

The cheapest monitor you can run

Log every field value the agent read at decision time, next to the decision it made. Not the prompt, the fields.

When something goes wrong six weeks later, you will want to know what the data said at the moment of the action, and no warehouse snapshot will reconstruct that. It also gives your evaluation work something real to run against. That is where a proper evaluation suite and production drift monitoring stop being theoretical.

The order that works, and what each step costs. Each step is bounded by the workflow named in step 1. Unbound, this becomes the programme that fails. 1 2 3 4 Scope One workflow. List the fields. One afternoon Repair Resolve entities, then set freshness. Weeks, not quarters Context One definition per term. Add lineage. Ongoing, small Gate Fail closed. Log field values. Never finished Step 4 is the only one with no end state, which is why it belongs to whoever operates the agent.
Read step 1 as the gate on the other three. Attempting steps 2 to 4 across an entire estate is the failure mode, not the fix.

Where this argument is weakest

Read the survey numbers above and the obvious conclusion is to stop shipping agents and go fix the data. That conclusion has a worse track record than the thing it replaces.

Gartner's 80% governance failure prediction is about precisely this class of programme. Data teams that started 18 month readiness efforts in 2024 largely now have neither ready data nor a shipped agent, and they have spent the budget that would have funded either. The sequence above only works because every step is bounded by a named workflow. Unbind it and you have rebuilt the programme that fails.

If I had to choose between shipping a narrow agent on imperfect data and waiting for a clean estate, I would ship, with the refusal behaviour from step 4 wired in before anything else. Imperfect data with a system that knows when to stop beats clean data that arrives in 2028.

What the survey base cannot tell you

The 75.9% figure comes from a survey run by a data infrastructure company, answered by self-selected data leaders and practitioners. People who choose to answer a data company's survey are more likely than average to believe data is the binding constraint. Nothing in it is randomised.

No question isolates data quality from the confounders sitting next to it. Those include unclear scope, absent change management, and the plain fact that some of these workflows were poor automation candidates regardless of their data. Treat the ranking as real and the magnitude as directional.

No published study takes matched deployments, varies data readiness while holding everything else constant, and measures the difference in outcome. Until one exists, "data quality is the top blocker" means data leaders rank it first when asked.

That is a report of belief, held by informed people, and it is not the same object as a measured cause. Anyone quoting these numbers as proof that data caused the failures is overreaching, including most of the vendors currently doing so. The honest version is narrower and still actionable: the people closest to these systems consistently name the same constraint, and the remediation is cheap enough that you do not need certainty to justify it.

Frequently asked questions

What is AI-ready data?

AI-ready data is data qualified for a specific AI use case rather than certified in general. Gartner describes it as data aligned to the use case, governed at the asset level, supported by pipelines with quality gates, described by live metadata, and continuously quality-assured. The important consequence is that the same dataset can be ready for one agent and unfit for another, so readiness is assessed per workflow.

Why do AI agents fail because of data quality?

Agents act on data instead of displaying it, and they chain steps, so errors compound. A field that is correct 99% of the time survives a dashboard, because a person notices when a number looks wrong. Across a 20 step agent workflow the chance every step reads a clean value falls to about 82%, and no human is watching any individual step.

How much enterprise AI failure is caused by data quality?

No study isolates it. What exists is ranking. 75.9% of more than 540 data leaders surveyed by The Modern Data Company in August 2026 named data quality and trust a top-three barrier to production agents. That ranked ahead of context and lineage at 63.5% and security at 61.7%. Gartner separately expects 60% of AI projects unsupported by AI-ready data to be abandoned through 2026. Both are beliefs and forecasts, not measured causes.

Should we fix our data before deploying AI agents?

Not as a separate programme. Gartner predicts 80% of data and analytics governance initiatives will fail by 2027, and horizontal readiness projects are the ones that stall. Scope the repair to one workflow instead. List the fields that workflow reads and writes, fix identity and freshness on those fields only, define the terms the agent will use, then gate the pipeline. Ship the narrow agent while you do it.

What is the difference between data quality and AI data readiness?

Data quality asks whether a field is accurate, complete and consistent, and that has one answer per field. AI readiness asks whether the data carries enough context for a system to act on it without a person checking the result, and that answer changes with every workflow. Readiness also covers unstructured sources, which are 70% to 90% of enterprise data and largely outside conventional quality tooling.

How do you know when data is ready for an AI agent?

Test it against the workflow, not against a standard. Take the list of fields the agent reads, and confirm three things about each one. The entity it describes resolves to a single record. The value carries a timestamp inside a stated staleness budget. The term it represents has one written definition the agent can reach. Anything failing those three should make the agent refuse rather than proceed.

Where to start this week

Two things, both achievable before Friday, neither requiring a purchase.

First, take the agent closest to production and write out its field list. Every field it reads, every field it writes, and one line each on what breaks if the value is wrong. Do not clean anything yet. The list is the deliverable, and in most organisations nobody has ever written it down, which is itself the finding.

Second, pick the single highest-consequence field on that list and answer two questions about it. Can you name the system of record? Can you say how old the value was when the agent last read it? If either answer takes more than a minute to find, you have located where step 2 begins, and you have located it for the price of an afternoon rather than a platform.

If you take one thing

Readiness is a property of a pairing between a dataset and a task, never a certificate you award an estate. That is also why the data advantage people describe as a moat is usually narrower than claimed, a distinction taken apart in the piece on data moats and workflow moats.

References

  1. The Modern Data Company, annual enterprise data leaders survey, reported by ITBrief, 21 August 2026. More than 540 responses across 66 countries. Used for the 75.9%, 63.5%, 61.7%, 25.9%, 19.5%, 57.3%, 23.5%, 33.8%, 8.4%, 21.7%, 65.1%, 39% and 10% figures. Vendor-run and self-selected, so directional on magnitude.
  2. Gartner, Lack of AI-Ready Data Puts AI Projects at Risk, 26 February 2025. Used for the 60% abandonment prediction, the definition of AI-ready data, and the 63% figure from a Q3 2024 survey of 248 data management leaders. The 60% figure is a prediction, not an observation.
  3. TechTarget, AI-ready data needs its own set of rules, experts say, 13 March 2026. Used for the Melody Chien quote, the Roxane Edjlali framing of metadata, the 32% readiness-process figure and the 70% to 90% unstructured data range.
  4. Gartner, Gartner Says Lack of Semantics Causes Inaccurate AI Agents and Wasted Spending, 11 May 2026. Used for the semantics forecast and the Rita Sallam framing. Both the 80% accuracy and 60% cost figures are "up to" forecasts for 2027 with no published methodology.
  5. Gartner, Gartner Predicts 80% of D&A Governance Initiatives Will Fail by 2027, 28 February 2024. Used for the governance failure prediction and the Saul Judah comment on prioritised business outcomes.
  6. CIO Dive, AI project failure rates are on the rise, reporting S&P Global Market Intelligence research across more than 1,000 respondents in North America and Europe. Used for the 42%, 17% and 46% figures and for the obstacles respondents named.
  7. TechTarget, What enterprises are getting wrong about AI data readiness, 9 July 2026. Used for the use-case-first framing and the practitioner argument against enterprise-wide preparation.

The weakest thing about this source base: the two largest numbers in this post come from a vendor-run survey and a Gartner prediction, and neither establishes causation. No study anywhere varies data readiness across matched deployments and measures the result, so every claim here about data causing failure is a claim about what informed practitioners report, not about what has been demonstrated. The 82% figure in the compounding section is arithmetic on an assumed error rate, not an observed value.

SK
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Works in People and Culture at Wayground (Quizizz), and writes here on what AI actually does to how software teams work, hire and are measured.

Related reading