From Aryan Vatsa | Product & Market Analysis

Harvey vs Legora vs General AI: Which Legal Tools Earn the Seat Price

On this page

Legal-specific research tools and a general assistant scored the same 80% on a 210-question legal research benchmark. The lawyer control group scored 71%. That result does not make Harvey or Legora overpriced, and it does change what you are buying when you pay for legal AI tools: workflow, provenance and an audit trail, not raw answer quality.

Key takeaways

  • Accuracy has stopped being the reason to buy legal-specific AI. In the Vals AI legal research benchmark, legal-specific tools and ChatGPT both averaged 80% against a 71% lawyer baseline across 210 questions.
  • The premium buys workflow, permissions and a defensible record. Harvey reports more than 25,000 custom agents running across 1,300-plus organisations. Legora passed $100 million in annual recurring revenue 18 months after launch.
  • Only one side of this comparison publishes a price. Claude Team lists $20 per seat per month billed annually and Microsoft 365 Copilot Business lists $18. Harvey and Legora publish nothing.
  • Legal-specific does not mean verified. A preregistered Stanford study found the purpose-built research tools from LexisNexis and Thomson Reuters hallucinated between 17% and 33% of the time.
80% vs 71%Legal-specific tools and ChatGPT both averaged 80% on legal research, against a 71% lawyer baseline. Source: Vals AI via LawSites, 2025.
69% vs 42%Legal professionals using general-purpose AI versus legal-specific applications, in the same survey. Source: 8am report via LawSites, March 2026.
$18 to $20Published seat price for a business-tier general assistant. Harvey and Legora publish no price at all. Source: Claude and Microsoft pricing pages, 2026.

The short answer for a legal buyer

Buy a legal-specific platform when the work product leaves your building: court filings, regulated advice, client deliverables and anything a regulator or opposing counsel can audit. Use a general assistant for internal drafting, summarising and first-pass research. The accuracy gap that once justified the premium has closed. The governance gap has not.

That is the whole decision compressed into five sentences, and most buyers get it wrong in the same direction. They run a bake-off on answer quality, watch both tools produce a competent memo, and conclude the expensive one is a wrapper. Answer quality is the one dimension where the two paths have converged.

The reader this post is written for is the person signing the contract: a general counsel, a managing partner or a head of legal operations who has to defend a line item. The tactical prompt-level material belongs elsewhere. The money question belongs here.

What Harvey and Legora are actually selling

Both companies describe themselves in the same terms now, which is itself informative. Neither positions on model quality. Both position on the shape of the work.

Harvey

Harvey announced a $200 million raise at an $11 billion valuation on 25 March 2026, co-led by GIC and Sequoia. Its own announcement puts the platform at more than 1,300 organisations and over 100,000 lawyers across 60 countries, including more than 500 in-house legal teams.

The figure that describes the product better than the valuation does is 25,000. That is the number of custom agents the company says now run on Harvey, across M&A, due diligence, contract drafting and document review. A customer is not buying a chat window with a legal system prompt. They are buying a place to put 200 firm-specific procedures.

The competitive backdrop matters when you read that valuation. Harvey sells directly against an incumbent with a much larger installed base, which is covered in the analysis of how Harvey stacks up against Thomson Reuters in legal AI.

Legora

Legora said on 28 August 2026 that it had passed $100 million in annual recurring revenue within 18 months of its October 2024 launch. It reports more than 1,000 customers across 50 markets, named clients including White & Case, Linklaters and Barclays, and more than 400 staff across nine offices. Its Series D closed at $600 million in April 2026 at a $5.6 billion post-money valuation.

Its stated product shift is identical to Harvey's. Usage moved from discrete tasks such as research or document review toward multi-step workflows that produce structured outputs and reports. In August 2026 it added a connection into NetDocuments so the agents can read a firm's own document management system.

That last detail is the actual product. Not the model, the connection. This is the pattern described in the piece on why integration depth behaves like a moat when the model does not.

The benchmark that complicates the premium

An independent evaluator, Vals AI, published a legal research benchmark in October 2025. It asked 210 questions across nine research types, including statutory definitions, 50-state surveys, multi-jurisdictional questions and sourcing recent caselaw.

Legal research accuracy, 210 questions, nine research types Weighted score. Accuracy 50%, authoritativeness 40%, appropriateness 10%. Lawyer baseline 71% Counsel Stack 81% Alexi 80% Midpage 79% ChatGPT 80% Blue bars are legal-specific tools. The gold bar is a general assistant. Grouped by category, both sets landed on the same 80% average.
The gap that matters here is vertical, not horizontal. Every tool beat the lawyers. No tool beat the general assistant by a margin a buyer could feel.

What the benchmark measured

Scoring weighted accuracy at 50%, authoritativeness at 40% and appropriateness at 10%. The authoritativeness weight is the interesting one, because that is exactly the dimension legal-specific tools are sold on. Even with 40% of the score pointed at whether an answer rests on real, citable authority, the generalist tied.

The report is explicit about its scope. It covers general legal research only, not drafting pleadings or producing formatted citations. Those are different tasks with different failure modes.

Where lawyers still beat the tools

Vals ran an earlier and broader study, the Vals Legal AI Report, across seven task types with a lawyer control group. Harvey led on data extraction, document question answering at 94.8% and transcript analysis, and tied the baseline on chronology generation. CoCounsel won summarisation at 77.2% against a 50.3% lawyer baseline.

The lawyer baseline beat every tool on two tasks: redlining and EDGAR research. Both are tasks where the value sits in judgement about what to change and what to look for, rather than in retrieving and compressing text. That split has held across every legal benchmark published so far, and it is the most useful thing in either report.

One caveat has to be said plainly, because it undercuts the headline. The October 2025 benchmark did not include Harvey or Legora. It compared smaller legal-specific research tools against ChatGPT. Nobody has published an independent head-to-head of the two most expensive platforms against a $20 assistant, and the absence of that test is a fact about the market, not an oversight.

Price is the one variable nobody publishes

Here is the asymmetry at the centre of this comparison. One side of it has a price page. The other side has a contact form.

The prices that are published

Anthropic lists Claude Team at $20 per seat per month billed annually, $25 monthly, with a premium seat at $100. Its Enterprise tier is described as seat price plus usage at API rates, and it routes you to a buying specialist. Microsoft lists Microsoft 365 Copilot Business at $18 per user per month paid yearly, promotional against a $21 standard price through 31 December 2026, and it requires a qualifying base licence underneath.

Those are real numbers you can read, budget against and compare. They are also the floor of the comparison, not the total: a base licence, a data processing agreement and an administrative burden sit under all of them.

What you can read on a price page, and what you cannot Published list price per seat per month, annual billing, read 29 August 2026 M365 Copilot Business $18 Claude Team $20 Claude Team premium $100 Harvey No published price. Quote only. Legora No published price. Quote only. The dashed bars have no length because no length is knowable. That is the point.
Two of these five rows can be checked by anyone. Three cannot, and the two most expensive products are in the group that cannot.

Why the circulating seat figures are not evidence

Search for Harvey or Legora pricing and confident numbers appear immediately. Ranges per seat, seat minimums, annual floors, renewal uplifts. Follow the citations and they resolve to comparison pages published by competing legal AI vendors, or to aggregators restating each other.

Six restatements of one unattributed figure is one source, not corroboration. So this post publishes no seat price for either vendor, and I would treat any number you find as marketing until the vendor itself states it. The correct reading of quote-only pricing is not that it is expensive. It is that price is being set per customer, which means your negotiating position is a bigger variable than the list price would have been.

That has a practical consequence at renewal, when the switching cost is highest and the comparison is hardest. The mechanics are laid out in the piece on what happens to a grandfathered price at the first real renewal.

Score the paths by use case, not by vendor

The vendor comparison is the wrong unit of analysis. A legal team does eight or nine distinguishable kinds of work, and the right answer differs by row. Most teams will end up running both paths, which is a defensible outcome rather than a failure to decide.

Three paths, what each is, and what is actually knowable about the price
PathWhat it isPublished priceWhat the money buys
General assistant, business tierChatGPT, Claude or Copilot on a business or enterprise agreement.Yes. $18 to $20 per seat per month, annual billing.Reasoning, drafting and summarising, with commercial data terms.
Legal-specific platformHarvey or Legora, with agents, vaults and document system connections.No. Quote only, no public page.Workflow, firm knowledge, permissions and a reviewable record.
Legal research incumbentWestlaw or Lexis with AI features on licensed content.Partly. Existing subscription plus AI modules.Licensed primary content and citator coverage.

The third row is included because most buyers already pay for it and forget to count it. Any honest cost comparison starts from what the team spends today, not from zero.

Which path wins, row by row Scored on fit for the task, not on model quality General assistant Legal platform Verdict Internal first drafts Strong Overkill General Summarising your files Strong Marginal General Research with citations Adequate Adequate Verify either Bulk contract review Weak Strong Platform Client and court output Weak Strong Platform Firmwide permissions Weak Strong Platform Two rows favour the cheap path outright. Three favour the platform. One favours neither.
Notice the shape. The platform wins where the output is governed, and loses where the output never leaves the drafter's screen.
The use-case scorecard, with the test that decides each row
Use caseBetter pathThe test that decides it
Internal memos and first draftsGeneral assistantWould anyone outside the team ever see this file?
Summarising documents you already holdGeneral assistantIs the source document already verified by a human?
Legal research with citationsEither, with mandatory verificationWho checks every citation, and is that checking logged?
Contract review at volumeLegal platformAre you reviewing more than a handful of agreements a week?
Client deliverables and court filingsLegal platformCan you reconstruct how the output was produced, months later?
Firmwide rollout with matter permissionsLegal platformDoes the tool need to respect an ethical wall on its own?

Where general AI genuinely fails

The cheap path has three real failure modes, and none of them is answer quality. Every one of them is about the record.

Confidentiality is a contract question, not a model question

The ABA published Formal Opinion 512 on 29 July 2024, its first ethics guidance on generative AI. It runs through competence under Model Rule 1.1, confidentiality under 1.6 and supervision under 5.3. It also states that the required verification depends on the specific tool and the specific task, which is a standard no vendor can satisfy on your behalf.

The operative distinction is not open model against legal model. It is consumer account against a business agreement with data terms you have actually read. A firm running client matter text through personal logins has a problem that no amount of vertical software fixes. It is the same exposure quantified in the analysis of what shadow AI adds to the cost of a breach.

If your team is already on general assistants, the governance work is the purchase. The comparison of enterprise governance across ChatGPT, Claude and Gemini is the right place to start on which terms differ.

The record that is actually accumulating

A public database maintained by the researcher Damien Charlotin tracks court decisions in which a party relied on AI-hallucinated material and a court responded. It stood at roughly 1,980 recorded decisions in late August 2026, across multiple jurisdictions, and it grows most weeks.

Read that number carefully, because the vendor version of it is misleading. A large share of those decisions involve self-represented litigants, and almost all of them involve output that nobody verified. The tracker records a verification failure, not a tool failure.

Legal-specific software does not exempt you from that record either. The Stanford RegLab study Hallucination-Free?, a preregistered evaluation across more than 200 legal queries, found that the purpose-built tools from LexisNexis and Thomson Reuters hallucinated between 17% and 33% of the time. Lexis+ AI answered 65% of queries accurately. Westlaw AI-Assisted Research answered 42% accurately.

Both vendors had marketed hallucination-free citations. That is the strongest available argument against paying a premium for the word "legal" on the tin. It is also why the only architecture that survives contact with a court is the one described in the piece on designing the human review step before the automation.

Where this argument is weakest

Four things in this post are softer than the framing implies, and a buyer should know all of them before quoting any of it.

First, the 80% headline does not include Harvey or Legora. Nobody has run the test this post is nominally about. The finding generalises to legal research, which is one row of a nine-row scorecard, and it may not generalise to agentic contract workflows at all.

Second, the price comparison is asymmetric by construction. Comparing a published list price against a blank space makes the gap look enormous, and a negotiated enterprise quote against 200 general assistant seats plus governance overhead may be closer than it appears.

Third, the Stanford data is from 2024 evaluations, peer reviewed in 2025. Every product tested has shipped substantially since. The finding that legal-specific tools hallucinate is durable. The specific percentages are not current, and should be read as a caution about vendor claims rather than a current scorecard.

Fourth, benchmarks measure tasks and procurement buys programmes. A tool that scores worse and gets used by 80% of the team beats a tool that scores better and gets used by nine people. No published benchmark measures that, and it is the variable that most often decides whether the spend returned anything. The wider version of that argument sits in the analysis of what actually makes an AI product defensible when the model is commodity.

Frequently asked questions

Is Harvey AI worth the price compared to ChatGPT?

It depends entirely on where the output goes. For internal drafting and summarising, published benchmark evidence shows a general assistant matches legal-specific research tools on accuracy at a fraction of a published price. For client deliverables, court filings, high-volume contract review and firmwide rollouts that must respect matter permissions, the platform buys workflow and an audit trail that a general assistant does not provide at any tier.

What is the difference between Legora and Harvey?

Both sell agentic legal workflows rather than a chat interface, and both are converging on the same description of the product. Harvey reported an $11 billion valuation in March 2026 with more than 1,300 organisations and over 25,000 custom agents. Legora passed $100 million in annual recurring revenue in August 2026 with more than 1,000 customers across 50 markets. Neither publishes pricing, so a direct cost comparison is not possible from public information.

Can I use ChatGPT for contract review?

For a handful of agreements reviewed by a lawyer who checks the output, yes. For volume, no. The failure is not comprehension, it is process: no clause library tied to your positions, no consistent redline history, no permissions model and no record of who reviewed what. Bulk contract review is the clearest row on the scorecard where a legal-specific platform earns its price.

Do legal AI tools still hallucinate?

Yes. A preregistered Stanford RegLab study found that purpose-built legal research tools from LexisNexis and Thomson Reuters hallucinated between 17% and 33% of the time, against vendor marketing that had claimed hallucination-free citations. Legal-specific grounding reduces the rate, it does not remove it. Every citation still needs a human check, and that check should be logged rather than assumed.

How much does Harvey AI cost per seat?

Harvey does not publish pricing, and neither does Legora. Seat figures circulating online trace back to aggregator pages and competing vendors rather than to either company, so none of them can be verified. What is knowable is the other side of the comparison: Claude Team lists $20 per seat per month on annual billing, and Microsoft 365 Copilot Business lists $18 per user per month with a base licence required.

Which legal AI tool is best for an in-house legal team?

In-house teams usually start in a different place from law firms, because they already hold enterprise licences for a general assistant. The practical sequence is to govern what the team is already using, measure which tasks it fails at, then buy a legal platform against those specific failures. Buying the platform first produces a large contract and a low adoption rate, which is the most common negative outcome in this category.

How to settle this in four weeks

The bake-off most teams run is the wrong experiment, because it tests the dimension that has already converged. Run this instead.

In weeks one and two, log every AI-assisted task your team already performs, with the tool used and whether the output left the building. Most legal teams discover that 70% of usage is internal drafting on tools they already pay for, which immediately prices the platform decision against a much smaller slice of work than the vendor deck assumes.

In weeks three and four, take the tasks that did leave the building. Ask one question of each candidate vendor: reconstruct, for a matter closed six months ago, which documents an agent saw and who approved the output. A platform that can answer that has sold you something a $20 seat cannot. A platform that cannot has sold you a chat window with better fonts, and the general assistant is the correct purchase.

Related analysis

This sits alongside the breakdown of Harvey against the legal research incumbent and the wider question of whether vertical AI is genuinely eating horizontal software.

References

  1. LawSites, Vals AI's latest benchmark finds legal and general AI now outperform lawyers in legal research accuracy, October 2025. Used for all 210-question benchmark figures and the 71% lawyer baseline.
  2. LawSites, AI adoption among legal professionals has more than doubled in a year, 5 March 2026. Used for the 69% and 42% adoption split, sample of more than 1,300 professionals.
  3. Legora, Legal teams' adoption of AI propels Legora past $100 million in ARR, 28 August 2026. Used for ARR, customer count and product description.
  4. Harvey, Harvey raises at $11 billion valuation, 25 March 2026. Used for valuation, organisation count and the 25,000 agents figure.
  5. Anthropic, Claude pricing, read 29 August 2026. Used for the $20 and $100 seat prices.
  6. Microsoft, Microsoft 365 Copilot Business, read 29 August 2026. Used for the $18 promotional price and the base licence requirement.
  7. Stanford RegLab, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools, 2024 preprint, published in the Journal of Empirical Legal Studies, 2025. Used for the 17% to 33% hallucination range.
  8. Damien Charlotin, AI Hallucination Cases database, read 29 August 2026. Used for the count of recorded decisions.

The weakest part of this source base is the price side. Harvey and Legora publish nothing, so this post deliberately quotes no seat figure for either, and the resulting comparison is one-sided by necessity. The hallucination case count changes daily and was read from the database's published summary rather than a full crawl.

AV
Aryan Vatsa
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading