From Mihir Katiyar | Product & Market Analysis

The Real AI Moat in 2026 Is Integration Depth, Not Model Depth

On this page

The top four AI labs finished March 2026 inside 25 Elo points of each other. That spread is narrower than the gap between two generations of the same product. If frontier capability no longer separates the labs, it cannot separate the products built on top of them. The AI moat that survives 2026 is integration depth, and this post makes the case with deployment evidence rather than assertion.

Key takeaways

  • Frontier capability converged during 2026. Anthropic, xAI, Google and OpenAI sat within 25 Elo points on the Arena leaderboard in March 2026, and the top US model led the top Chinese model by 2.7%.
  • The price of a fixed capability falls about 40x a year. Epoch AI measured the cost of GPT-4 level performance on PhD-level science questions falling from $37.50 per million tokens in March 2023 to $0.12 by December 2024.
  • Enterprises already buy the model as a component. Three vendors take 88% of enterprise LLM API usage, and Microsoft reported a fivefold rise in 2026 in customers building against more than one provider.
  • The deployments that stuck are wired into systems of record. Abridge is live in more than 300 health systems through Epic, Oracle Health and athenahealth. Harvey sells the connection to iManage, NetDocuments and SharePoint.
25 EloSpread separating the top four labs on the Arena leaderboard, March 2026. Source: Stanford AI Index 2026.
40x a yearFall in the price of GPT-4 level science reasoning per million tokens. Source: Epoch AI.
88%Share of enterprise LLM API usage held by three model vendors. Source: Menlo Ventures, 2025.

What integration depth actually means

Integration depth is how far a product reaches into the systems a customer already runs. It is also how much work it writes back into those systems under the customer's own identity, permission and audit rules. Model depth is the capability you would lose by swapping the underlying model. In 2026 the first is scarce and the second is not.

The distinction is practical, not philosophical. Model depth is bought on a price list that every competitor can read. Integration depth is earned one customer system at a time, and it does not transfer when a rival raises a round.

The plumbing decides the deal. The model decides the demo.

Integration depth is not an integrations page

A logo grid of 200 connectors is a marketing asset. Depth is a different measurement, and it is countable.

Count how many of those connections write rather than read. Count how many respect the customer's own role model instead of asking for a service account with wide access. Count how many survive a permission change halfway through a task. Most products score close to zero on all three, which is why most products get replaced quietly.

The model layer converged, then it got cheap

Two measured trends sit underneath this argument. Capability at the frontier is clustering, and the price of any given capability level is collapsing.

Four labs, 25 Elo points apart

The 2026 AI Index from Stanford HAI records four companies clustered within 25 Elo points on the Arena leaderboard as of March 2026. Anthropic sat at 1,503, xAI at 1,495, Google at 1,494 and OpenAI at 1,481.

The same chapter puts the top US model ahead of the top Chinese model by 2.7% as of March 2026, down from a double-digit gap in 2023. The index is explicit that competition has moved toward cost and reliability rather than raw capability.

One number cuts the other way and deserves stating. The top closed model leads the top open model by 3.3%, up from 0.5% in August 2024. Open weights did not catch the frontier in 2026, they fell slightly further behind it. Convergence is happening among the closed labs, not across the whole field.

The frontier is a cluster, not a ladder Arena leaderboard scores, March 2026. Axis runs 1,470 to 1,510 only. 1,470 1,490 1,510 Anthropic1,503 xAI1,495 Google1,494 OpenAI1,481 22 points, first to fourth
Notice the axis. It spans 40 points, because a full-width axis would show four dots on top of each other.

Price is the second half of the story. Epoch AI tracked what it costs to buy a fixed capability level over three years and found declines ranging from 9x to 900x per year depending on the benchmark. Reaching GPT-4 level performance on GPQA Diamond cost $37.50 per million tokens in March 2023 and $0.12 by December 2024.

Epoch is careful about the caveat, and so should you be. The fastest declines happened in the most recent year, which is exactly the period least likely to repeat. What that means for gross margin is a separate argument, covered in the piece on inference costs and AI margins.

What a fixed capability level costs, before and after Price per million tokens to hit the same benchmark score. Vertical scale is logarithmic. First model at that level Cheapest model, later $37.50 $2.00 $0.12 GPQA $0.10 HumanEval $0.07 MMLU March 2023 2024 prices Source: Epoch AI, LLM inference price trends. Benchmarks: GPQA Diamond, HumanEval, MMLU.
The capability did not get better in this chart. Only the price moved, and it moved by more than two orders of magnitude.

Where enterprise deployments actually fail

If models were the constraint, the failure pattern would look like wrong answers. It does not. It looks like software that works in a demo and dies on contact with a real workflow.

The learning gap, and what that study cannot prove

MIT's NANDA project reported in 2025 that the large majority of enterprise generative AI pilots produced no measurable return. It attributed the shortfall to what it named a learning gap. The description was about workflow fit and retained context, not model error.

Be honest about the evidence. That work rests on 52 executive interviews, 153 survey responses and 300 public deployments, it is not peer reviewed, and its headline percentage has been quoted far past what the sample supports. The direction is credible. The precision is not. The return question itself is handled separately in the analysis of where measurable AI return has shown up.

Menlo Ventures found the commercial signature of the same problem from the other side. AI buyers convert at 47% against 25% for traditional software, because the value is legible in the first meeting. Fast conversion and weak returns is the pattern of a market that buys capability quickly and cannot install it.

Three products whose moat is plumbing

Assertions about moats are cheap. Deployment counts are not. Each of these companies sells into a market where every competitor can reach the same models.

Abridge: 300 health systems, three EHRs

Abridge said in its 11 June 2026 platform announcement that it is live in more than 300 health systems. It also reports over 100 million conversations a year, and partner systems covering more than 250 million patients. It integrates with Epic, Oracle Health and athenahealth.

None of that is a model claim. A rival with an equal or better speech model still has to earn a place inside the electronic health record. Then comes the security review, the note-to-billing fit, and a clinical governance committee. That sequence is measured in quarters.

Harvey: the document system is the product surface

Harvey's own argument to mid-sized law firms is not about reasoning quality. It is that lawyers otherwise spend their day moving files out of a document management system and into an AI tool by hand. So the product connects to iManage, NetDocuments, SharePoint, Box and Microsoft 365, and adds licensed content through the LexisNexis alliance.

Read that as a positioning choice. A company at Harvey's valuation could compete on legal reasoning benchmarks and chooses to compete on where the documents live. The content partnership angle is examined further in the piece on legal AI and incumbent content owners.

Glean: the horizontal version of the same bet

Glean publishes a count of 275 or more connectors as its headline product fact. For a horizontal platform this is the only available moat, because horizontal buyers can switch assistants in an afternoon. The strategic split between the two shapes is the subject of the analysis of vertical AI eating horizontal SaaS.

What each company sells access to, and what a rival must rebuild
CompanySystems it is wired intoWhat a well-funded rival must reproduce
AbridgeEpic, Oracle Health, athenahealth, across 300+ health systemsClinical governance approval, note-to-billing fit, security review per system
HarveyiManage, NetDocuments, SharePoint, Box, Microsoft 365, LexisNexis contentMatter-level permissions, versioning and metadata on export, licensed content rights
Glean275+ enterprise application connectorsPer-source permission mirroring, freshness guarantees, connector maintenance

Deployment and connector counts in this table are company-reported and not independently audited. The right-hand column is my judgement about replication cost, not a measured figure, and a rival with a distribution advantage may need none of it.

Microsoft is telling you the model is the swappable part

The most useful evidence for this argument comes from the company with the most to lose from it. On its fiscal Q4 2026 earnings call on 29 July 2026, Satya Nadella told enterprises to keep the model separate from the harness. That harness holds their memory, context and action space, which is what makes any model swappable.

That is an architecture recommendation from the largest distributor of enterprise AI, and it describes the moat precisely. The harness is the moat. Microsoft reported a fivefold increase during 2026 in customers building with models from more than one provider, and lists over 11,000 models in its catalogue.

I read this as strategy, not neutral advice. Microsoft benefits enormously from a world where models are interchangeable and the harness sits in Azure. That does not make the description wrong. It makes it a description written by someone who intends to be on the winning side of it.

What integration depth is actually made of

The word integration hides six separate jobs, and they are not equally hard. Most teams do the first three and call the work finished.

Reading data is close to free now. Writing into a system of record under someone else's permission model, with an audit trail a regulator will accept, is where the cost sits. That asymmetry is the whole moat.

Six layers, three levels of difficulty A judgement grid, not measured data. Read it as where effort accumulates. LAYERWHAT BLOCKS A COPYCOPY COST Model accessNothing, it is a price listLow Prompting and evalsWeeks of iteration, then parityLow Read-only connectorsMaintenance, not noveltyMedium Write-back to the recordError states, rollback, liabili…High Permission and auditCustomer's own role model, per…Highest Change managementHabits, training, internal poli…Highest The two layers a rival cannot buy are the two most teams never build.
The bottom two rows are not engineering problems. They are the reason a switch takes two quarters instead of two weeks.
The six layers, and who owns the difficulty
LayerWhere the work livesWhy it is or is not defensible
Model accessAn API keyEvery competitor has the same one, at a price falling yearly
Prompting and evaluationYour teamReal skill, reproducible in a quarter by a good team
Read-only connectorsYour engineering backlogTedious, increasingly commodity, still necessary
Write-back to the recordShared with the customerCarries liability, so it is negotiated once and rarely reopened
Permission, audit and residencyThe customer's security teamBespoke per tenant, and the first thing a switch would break
Change managementThe customer's operationNot transferable, and the reason incumbency compounds

The interoperability standard question sits underneath all of this. The Model Context Protocol makes tool calls portable, which lowers the cost of the top three rows. It does not decide who may write what into a system of record, and that is the part that is expensive.

Where this argument is weakest

Three objections are strong enough that I would not publish this without them.

Integration depth decays

The MCP project's 2026 roadmap, published 9 March 2026, names enterprise readiness as a priority area. Its list is exactly what I have called the moat: audit trails, single sign-on integrated auth, gateway behaviour and configuration portability. If a working group standardises those, the barrier drops for everyone.

That is the honest counter-case. A moat made of undone work erodes as the work gets done elsewhere. It buys years, not permanence, and anyone selling it as permanent is overclaiming.

Sometimes the model does decide

Menlo's data shows Anthropic holding an estimated 54% of the coding segment against 21% for OpenAI. In that category the model is the product, the capability gap is visible in daily use, and buyers move to it. Any claim that model choice never matters dies on that number, and the buy-side version of the question is worked through in the build versus buy analysis for coding agents.

The narrower claim survives. Where a task has a hard capability ceiling, model depth decides. Where it does not, and that is most enterprise work, integration decides.

The evidence here is thinner than it looks

Every deployment figure in this post is self-reported by a vendor with an interest in the number. I cannot separate Abridge's integration advantage from its sales force, its brand or its funding, and neither can anyone else from public data. What I have is a pattern across three companies in three markets, which is suggestive and not conclusive. The related argument about whether thin products deserve the criticism they get is handled in the piece on the wrapper insult.

How to score your own integration depth

The framing above is only useful if it changes what you measure on Monday. Score your product out of 10 using the weights below, and score your closest competitor at the same time.

A weighted test you can run in an afternoon
CriterionWeightWhat a weak answer sounds like
Systems you write to, not just read from3"We can export to CSV"
Runs under the customer's own permission model3"We use a service account"
Audit trail a compliance team has accepted in writing2"We log everything"
Time from contract to production, last three customers1"It varies"
Model swap would be invisible to the customer1"We are built on the best model"

The last row is deliberately inverted. If swapping your model would be visible to a customer, you are carrying model risk you do not control. If it would be invisible, your value has already moved into the harness.

Depth in one customer's stack is not depth in the category, and a product embedded in the workflows of a shrinking buyer is embedded in a shrinking market. The wider consolidation pressure is covered in the piece on SaaS sprawl and rationalisation.

Frequently asked questions

What is the real moat for AI startups in 2026?

Integration depth, meaning the systems your product reads from and writes back into under the customer's own identity and audit rules. Frontier model capability has converged, with the top four labs inside 25 Elo points of each other in March 2026 on the Arena leaderboard. A capability every competitor can buy at falling prices is an input, not a moat. What a customer cannot rip out in a week is.

Is integration depth actually defensible or just hard work?

It is hard work that compounds, which is what a moat is. Each connection carries permission mapping, error handling, audit trails and a change management cost inside the customer. A rival must rebuild all of it before a buyer will switch, and the buyer must absorb the disruption. That is a real barrier. It is also a decaying one, because standards and tooling keep lowering the cost of the first connection.

Does the Model Context Protocol remove the integration moat?

Not yet, and the protocol's own roadmap explains why. The 2026 roadmap published in March lists enterprise readiness as a priority area still to be delivered, naming audit trails, single sign-on integrated auth, gateway behaviour and configuration portability. Those are the expensive parts of an enterprise integration. A standard that makes tool calls portable does not decide who may write what into a system of record.

How many foundation model vendors do enterprises actually use?

Menlo Ventures surveyed 495 US enterprise AI decision makers in November 2025 and found three vendors accounting for 88% of enterprise LLM API usage. Anthropic held 40%, OpenAI 27% and Google 21%. Microsoft reported a fivefold increase during 2026 in customers building with models from more than one provider. Buying from several vendors is now normal, which is another way of saying the model is a component.

Why do enterprise AI pilots fail when the models are good?

Because the failure is rarely in the model. MIT's NANDA project reported in 2025 that most enterprise pilots produced no measurable return. It blamed a learning gap, meaning tools that do not fit existing workflows and do not retain context. That study rests on 52 interviews, 153 survey responses and 300 public deployments, and it is not peer reviewed. Treat the direction as sound and the precision as weak.

How do I measure integration depth when evaluating a vendor?

Ask three questions with numeric answers. How many customer systems does the product write to, not just read from. What happens to that written record when a user's permissions change mid task. How long did the last comparable customer take from contract to production use. A vendor with real depth answers all three from memory. A vendor without it answers with a logo wall.

Where to start this week

Start with a removal test. Pick your largest account and write down, in order, what would break if your product were switched off on Friday. If the list is short, or reads as inconvenience rather than breakage, you have a usage number and not a moat.

Then move one read-only connection to a write. One is enough to learn the real cost, because the permission mapping and the rollback behaviour will take longer than the API call by an order of magnitude. That gap is the moat, measured in your own engineering hours.

If you buy rather than build, invert the same test on your vendors and ask what your exit costs.

The buyer's side of this argument

Integration depth is a moat for the vendor and a switching cost for you. Both readings are correct, which is why the removal test above is worth running from whichever side of the contract you sit on.

References

  1. Stanford HAI, 2026 AI Index Report, Technical Performance. Used for Arena leaderboard scores, the 25 Elo cluster, the 2.7% US to China gap and the 3.3% closed to open gap.
  2. Epoch AI, LLM inference prices have fallen rapidly but unequally across tasks. Used for the 40x annual decline and the benchmark price pairs.
  3. Menlo Ventures, 2025: The State of Generative AI in the Enterprise, December 2025. Survey of 495 US decision makers, 7 to 25 November 2025. Used for vendor share, the 88% concentration, conversion rates and the coding segment split.
  4. Abridge, platform announcement, 11 June 2026. Used for health system count, conversation volume and EHR integrations.
  5. Harvey, Why an integrated AI platform matters for mid-sized law firms. Used for named document system integrations and the stated positioning.
  6. Model Context Protocol, The 2026 MCP Roadmap, 9 March 2026. Used for the enterprise readiness priorities quoted in the limitation section.
  7. Silicon Canals, Nadella pitches swappable AI models, 30 July 2026, reporting Microsoft's fiscal Q4 2026 earnings call. Used for the harness comment, the fivefold multi-provider figure and the catalogue size.

The weakest thing about this source base: every deployment figure is vendor self-reported and none has been independently audited. The MIT NANDA finding referenced in the failure section is preliminary, not peer reviewed, and is quoted here for direction only. The copy-difficulty grid is my judgement, not measured data.

SK
Mihir Katiyar
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading