From Mihir Katiyar | Product & Market Analysis

Data Moats Are Overrated. Workflow Moats Decide Who Survives

On this page

Every AI pitch deck claims a proprietary data advantage. Most of them describe an asset a competitor can buy, scrape or approximate within a year. The advantage that actually holds sits one layer up, in the workflow where a decision gets made, because that is the thing a customer would have to rebuild in order to leave.

Key takeaways

  • Most proprietary data carries a price. Reddit reported $43 million of other revenue in Q2 2026, predominantly data licensing, up 24% year over year, with OpenAI and Google named as the two largest buyers. An asset with a rate card is an input cost.
  • The data advantage that held in court held on copyright, not on volume. Thomson Reuters won against Ross Intelligence over 2,243 Westlaw headnotes, and the protected thing was editorial work product, not an accumulated corpus.
  • Workflow ownership shows up in contracted revenue, not in storage. ServiceNow reported $13.20 billion of current remaining performance obligations and 658 customers above $5 million in annual contract value in Q2 2026.
  • The only test that discriminates is what a leaver has to rebuild. A data advantage costs a departing customer one export. A workflow advantage costs them their approval chain, their audit trail and six months of retraining.
149%Palantir's US commercial revenue growth in Q2 2026, from a company that sells the decision workflow rather than the data.
2,243Westlaw headnotes a US court found Ross Intelligence had infringed, in the first ruling to reject fair use for AI training.
$43MReddit's quarterly data licensing revenue, which is what a data moat looks like once it has a published price.

Where the data moat claim came from

The argument runs like this. Your product collects data nobody else has. That data trains a model that works better than a competitor's. Better output attracts more users, who generate more data, and the gap widens on its own.

It is a clean story and it sells rounds. It is also, in the version most decks present, a description of a mechanism nobody has measured in that company.

The objection is 7 years old

Martin Casado and Peter Lauten set this out at Andreessen Horowitz in May 2019, well before the current cycle. Their central distinction is between a data network effect and a data scale effect, and almost every deck claiming the first is describing the second.

A network effect requires nodes interacting over a defined interface. More data improving a recommendation is not that. It is a scale effect, and scale effects behave badly at the margin.

Their sharpest observation is about cost. In a normal economy of scale each additional unit gets cheaper. With data the opposite holds: the cost of adding genuinely new records rises while the value of each incremental record falls. Their worked example put the useful ceiling well below full coverage, after which more collection bought nothing.

Nothing in the last 7 years has reversed that curve. What changed is that the floor moved.

What foundation models changed

Before 2020 a company starting in a narrow domain had to build its own corpus from zero. That bootstrap was expensive enough to look like a barrier. A pretrained model now arrives already knowing most of the domain, which means the minimum viable corpus a challenger needs has collapsed.

The second change runs the other way and is more interesting. Labs have worked through the readily available public text. Epoch AI's 2024 analysis put the effective stock of public human text at roughly 300 trillion tokens and had frontier training runs meeting that stock somewhere between 2026 and 2032.

Scarcity turned data into a market. Once something has a market, it has a clearing price, and a clearing price is the opposite of a moat. Your rival does not need to accumulate what you accumulated. They need a budget.

Three tests a data moat has to pass

Run any claimed data advantage through these before believing it. Failing one is common. Failing all three means the deck is describing a database.

Can a competitor buy it?

Reddit's other revenue line, which is mostly data licensing, reached $43 million in Q2 2026, up 24% year over year, with OpenAI and Google as the two largest buyers. The reported Google arrangement runs at roughly $60 million a year and comes up for renewal.

Read that as a buyer rather than as a shareholder. One of the largest user-generated text corpora in existence rents for a sum that a well funded competitor can approve in a single meeting. Almost every corpus below that scale is cheaper still.

My position is that a data asset with a rate card should be modelled as a cost line, not as a barrier. It affects your margin. It does not affect whether someone can enter.

Can a model approximate it?

The second test is harder to run and matters more. Take the task your data supposedly wins, and measure a general model against your specialised one on the same evaluation set.

If a general model lands within a few points, your corpus bought a lead, not a barrier, and the next model release will spend some of that lead for you. The gap has to be both large and stable across releases before it counts. Most teams have never run this comparison, which is why the claim survives.

The uncomfortable version of the finding: the more your data resembles text that exists elsewhere in the world, the more cheaply it is approximated. Data that is genuinely absent from the public record is the only kind that resists this, and there is much less of it than founders assume.

Does using the product create more of it?

The third test separates a compounding asset from a static one. When a customer uses your product today, does the product get measurably better tomorrow, in a way the customer would notice if it stopped?

A one-off data migration is a static asset. Every record was created once and never renews. A correction loop, where a professional overrides an output and the override becomes training signal, is a compounding asset, and it is rare because it requires the product to sit where the correction happens.

Notice what that requirement implies. The compounding version of a data advantage is a consequence of workflow position, which is the argument of this piece arriving from the other direction.

What a competitor has to reproduce, and how long it takes Directional ranking of four advantages by time to replicate. Not measured data. Access to a frontier model Days A licensed or scraped corpus Months Usage data your product creates Years The workflow a decision runs through Longer, if at all Model access is the shortest bar because every competitor buys it on the same commercial terms. Assessment based on observed replication patterns across AI application categories, not a measured study.
The two bars at the top are what most decks call a moat. The bar at the bottom is the one that survives a competitor with funding.

What the Westlaw ruling actually proves

There is one recent case where a data advantage held decisively, and it is worth reading closely because it does not prove what people cite it for.

In Thomson Reuters v. Ross Intelligence, decided in Delaware on 11 February 2025, the court granted partial summary judgment for Thomson Reuters and found that Ross had infringed 2,243 Westlaw headnotes. It was the first US ruling to reject a fair use defence for AI training material.

The reasoning matters. The court treated the use as commercial and found it lacked a further purpose, because Ross was building a tool that competed directly with Westlaw. Ross had tried to license the content, been refused, and obtained it through a third party instead.

So the thing that held was not the size of the corpus. It was that Westlaw headnotes are original editorial work, authored by people, protected by copyright, and owned outright by the company selling access to them. Strip any one of those four properties and the case changes.

What it does not generalise to

Very little enterprise data has those properties. Transaction records, sensor readings, prices and dates are facts, and copyright does not protect facts. Customer data is usually held under a contract that gives the vendor a processing right, not ownership, which is why the data-rights clause is the first thing a competent buyer's counsel rewrites.

The practical read is narrow. If your data advantage is a body of original authored work that you own and can refuse to license, you have something real and enforceable. If it is a warehouse of customer facts, the Ross ruling offers you nothing. The legal analysis in the Harvey and Thomson Reuters comparison sets out how differently the two positions behave in the same market.

Why workflow ownership holds when data does not

A workflow moat is unglamorous. It is the claim that the sequence of steps by which a decision gets made, approved, recorded and audited runs through your product, and that moving it costs more than the software is worth.

The unit of lock-in is the decision, not the record

Data is portable by construction. Every serious enterprise contract now carries an export clause, and every buyer's procurement team checks it. Assume your customer can take their data out in an afternoon, because increasingly they can.

A workflow is not portable, because most of it is not in your product. It is in who approves what, at which threshold, with which evidence attached, and in the fact that 400 people learned it. That knowledge sits in the organisation, and it is expensive to rewrite even when the software is free.

This is also why interoperability standards do not dissolve the position. Shared connectors make it cheaper to read the data, as the analysis of the MCP interoperability standard sets out, and they do nothing about who holds the approval step.

What it looks like in reported numbers

Two public companies make the shape visible. Palantir reported Q2 2026 revenue of $1.935 billion, up 93% year over year, with US commercial revenue of $764 million, up 149%, and 73 deals of at least $10 million closed in the quarter. Palantir does not sell a dataset. It sells the layer where operational decisions get made.

ServiceNow reported Q2 2026 subscription revenue of $3.877 billion, up 24.5%, with current remaining performance obligations of $13.20 billion and 658 customers above $5 million in annual contract value. Contracted future revenue is the cleanest available proxy for lock-in, because it is money customers have already committed under contract.

State the weakness in that evidence immediately: growth is not durability. Both companies are also riding a spending wave, and neither number isolates the workflow effect from the demand cycle. What the figures do show is that the largest contracts in enterprise software are being written for the decision layer, not for the data layer.

Five questions, asked of both kinds of moat The same test applied to a data advantage and to a workflow position Data moat Workflow moat What decides it Can a competitor buy it? Often No Rate cards exist Can a model approximate it? Frequently Not the process Models generalise Does use create more of it? Sometimes By definition Usage is the asset What must a leaver rebuild? A file export The whole process Switching cost Who checks the claim? Nobody The renewal Evidence, not words
The last row is the practical one. A data claim is asserted in a deck and never tested. A workflow claim is tested every renewal cycle, by the customer, with money.

The data moats that are real

The consensus is not wrong so much as undiscriminating. Some data advantages are formidable, and it is worth being precise about which, because the properties do not transfer between them.

Four asset types, and how each holds up under the three tests above.

Four kinds of claimed data advantage, tested for purchasability, approximation and compounding.
AssetCan a rival buy it?Can a model approximate it?Verdict
Original authored work you own, such as editorial headnotes or a proprietary taxonomyOnly if you agree to sell itCopying it is an infringement risk, as Ross foundA real moat, enforced by copyright
Regulated data you are licensed to hold, such as clinical or credit recordsNo, the barrier is the licence and the consent chainNo, the constraint is legal rather than technicalA real moat, enforced by regulation
Outcome data your product alone observes, such as which recommendation was accepted and what followedNot directly, it has to be generatedOnly where similar behaviour is publicReal while you hold the workflow that makes it
A large corpus of general documents, transcripts or logsYes, licensing markets exist and clearUsually, this is what pretraining already coversAn input cost presented as a moat

Read the third row again, because it is where the two arguments meet. Outcome data is the strongest of the four for a startup, and it is not an independent moat at all. It is a dividend paid by workflow position, and it stops accruing the day the workflow moves.

The first two rows are genuinely defensible and almost never available to a new entrant. That is the honest asymmetry: the data moats that work are mostly held by incumbents who acquired them under conditions that no longer exist.

Workflow moat or integration, and how to tell

The failure mode is claiming a workflow moat for what is really a connector. Both look like depth on a slide. They behave completely differently under attack.

Three questions, answerable in an afternoon with your own account data.

First, does the product hold a step that a human is accountable for? An approval, a sign-off, a filing, a submission. If the answer is that it produces a draft somebody else approves elsewhere, the accountable step is not yours.

Second, when a customer churns, what breaks for someone other than the buyer? If only the person who signed the contract notices, you had a subscription, not a position.

Third, what is the median time from the churn decision to the last day of usage? A workflow position shows up as a long, expensive unwind. A tool shows up as a switch flipped at the end of the month.

An integration is not a workflow

Integration depth was a credible moat when every connector was bespoke and expensive. That was the era when counting integrations was a reasonable proxy for switching cost.

Standardised tool interfaces changed the economics of exactly that layer. When connecting to a system of record becomes a configuration step rather than a quarter of engineering, the count of integrations stops separating anybody from anybody.

What survives is the part that was never really about connectivity: the decisions that run through the product and the people accountable for them. The same distinction explains why per-seat pricing is under pressure while outcome-linked contracts are not, which is set out in the analysis of seat compression. It also explains why the category winners in vertical AI tend to be the products that took over a regulated process, not the ones with the largest training set.

What a departing customer has to rebuild, by layer, and roughly how long each takes.
LayerWhat a leaver has to rebuildTypical unwind
Model accessNothing, the replacement calls the same providersImmediate
Stored dataAn export and an importDays
IntegrationsReconnect the same systems through a standard interfaceWeeks
Reports and dashboardsRebuild the definitions and reconcile the historyA quarter
Approvals, controls and audit trailRewrite the process, retrain the staff, re-certify the controlsMultiple quarters

The unwind column is the moat, measured in the only unit that matters to the person deciding whether to leave. It also tells you where to invest: the bottom row is the one nobody buys their way past.

Where this argument is weakest

Four honest problems with the case I have just made.

The first is that I have used growth figures as evidence for durability, and they are not the same thing. Palantir and ServiceNow are growing into an unusually strong spending cycle. Neither figure separates a workflow effect from a demand effect, and the correct test would be retention through a downturn, which has not happened yet for this cohort.

The second is that workflow moats erode from above rather than from the side. An agent that operates across several systems does not have to displace your approval step. It can sit over it and turn your product into a called service, which is the absorption mechanism described in the analysis of point-product absorption. My argument understates that risk because it is early and the evidence is thin.

The third is that I have set the bar for a real data moat high enough that the largest counter-examples fall outside it. Search query logs, clinical record estates and payment networks are enormous data advantages that have compounded for years. The honest framing is not that data moats are fictional. It is that they are rarely available to the company claiming one in a Series A deck.

The fourth is a limitation of this publication rather than of the argument. This post has no first-party data behind it. It is a framework tested against public filings and one court ruling. That is a weaker form of evidence than a measured study of churn across a portfolio, so it should be read as an argument rather than as a finding. The related question of what the label "wrapper" does and does not settle is taken up in the defence of AI application companies.

Frequently asked questions

What is a data moat?

A data moat is the claim that a company holds data competitors cannot obtain, and that this data makes its product better in a way rivals cannot match. The strong version adds a loop: more usage produces more data, which improves the product, which attracts more usage. Most claims describe a large dataset rather than that loop, and the two behave very differently.

Are data moats real in AI?

Some are. Original authored work you own outright and regulated data you are licensed to hold both create genuine barriers, the first enforced by copyright and the second by law. General corpora of documents, transcripts or logs mostly do not, because licensing markets exist and clear at prices a funded competitor can pay. Reddit's data licensing revenue reached $43 million in a single quarter.

What is a workflow moat?

A workflow moat means the sequence by which a decision is made, approved, recorded and audited runs through your product. Its strength is not what you store, it is what a departing customer has to rebuild: the approval chain, the control evidence, and the working knowledge of the staff who use it. That rebuild is slow, expensive and mostly not portable.

How do you tell a real moat from a pitch deck claim?

Ask what a competitor with funding would have to do to reproduce it. If the answer is buying a licence or paying for labelling, that is a cost line rather than a barrier. Then ask what breaks for someone other than the buyer when a customer churns. If nobody outside the signing manager notices, the product is a subscription, not a position.

Can a foundation model replicate proprietary enterprise data?

It cannot reproduce the records themselves, but it frequently reproduces the capability the records were meant to deliver. The practical test is measuring a general model against your specialised one on the same evaluation set. If the gap is small, the corpus bought a lead that the next model release will partly spend. Only data genuinely absent from the public record resists this.

Where to start this week

Take your own defensibility slide and run the three tests on it in writing. Not as a discussion, as a document with numbers in it, because the claim survives precisely because nobody makes it specific.

Write down what a funded competitor would pay to obtain the equivalent of your data. If you cannot name a price, ask a data broker for one, and expect the answer to be lower than you would like.

Then pull your last 8 churned accounts and record how many days passed between the decision to leave and the final day of usage. That number is your workflow position, expressed in the unit your customers actually experience. If it is short, you know what to build next, and it is not a larger dataset.

References

  1. Palantir Technologies, Q2 2026 results, exhibit 99.1, 2 August 2026. Used for total revenue, US commercial revenue growth and deal counts.
  2. ServiceNow, Second quarter 2026 financial results, 22 July 2026. Used for subscription revenue, current remaining performance obligations and the count of customers above $5 million ACV.
  3. Reddit Inc, Second quarter 2026 results, 30 July 2026. Used for the other revenue figure and its year-over-year growth.
  4. Martin Casado and Peter Lauten, Andreessen Horowitz, The Empty Promise of Data Moats, 9 May 2019. Used for the distinction between data network effects and data scale effects and for the cost curve argument.
  5. Jenner & Block, client alert on Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence Inc., February 2025, reporting No. 1:20-cv-613-SB (D. Del., 11 February 2025). Used for the headnote count and the fair use holding.
  6. Pablo Villalobos and colleagues, Epoch AI, Will we run out of data? Limits of LLM scaling based on human-generated data, 6 June 2024. Used for the stock of public human text and the range over which training runs meet it.

The weakest part of this source base is the court ruling, reported here through a law firm summary rather than the docket itself. The Epoch AI range is next, because it is a conditional model output rather than an observation. The comparison between data and workflow moats is an argument, not a measurement, and no first-party churn data sits behind it.

AV
Mihir Katiyar
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading