From Aryan Vatsa | Product & Market Analysis

Deep Research Tools Compared: Which One Actually Hallucinates Least

On this page

The best deep research agent on citation accuracy scores 93.68%. The worst of the major five scores 77.96%. That spread means roughly one statement in five carries a source that does not support it. The agent that retrieves the most citations per report is also the one whose links most often lead nowhere. Nobody is selling you the number that matters.

Key takeaways

  • Citation accuracy across the five major deep research agents ranges from 77.96% to 93.68%. Claude with search leads at 93.68% and Perplexity Deep Research follows at 90.24%, measured across 100 expert-written tasks in DeepResearch Bench.
  • Volume and precision trade against each other, and vendors market volume. Gemini 2.5 Pro Deep Research returns 111.21 supported citations per report, three times the next agent, while scoring fourth of five on accuracy.
  • 13.3% of the URLs in Gemini 2.5 Pro Deep Research reports resolve nowhere and have no archive snapshot. The equivalent figure for OpenAI Deep Research is 3.5% and for Claude with search is 3.2%.
  • Search grounding is the single intervention that matters most. The same model families hallucinate 21% to 59% of citations when generating unassisted, against 3% to 5% once retrieval is attached.
93.68%Highest measured citation accuracy, Claude-3.7-Sonnet with search, against 77.96% for OpenAI Deep Research. Source: DeepResearch Bench, 100 tasks.
13.3%Share of Gemini 2.5 Pro Deep Research URLs that resolve nowhere and were never archived. Source: arXiv 2604.03173, 53,090 URLs.
60%Share of 1,600 citation queries that eight AI search tools answered incorrectly. Source: Tow Center, March 2025.

What "hallucination" actually means in a research report

A deep research agent runs a long chain of searches, reads what it finds, and returns a report with inline citations attached to specific claims. The failure everyone imagines is the fabricated paper, complete with plausible authors and a DOI that goes nowhere. That failure is real, and it is not the common one.

Three distinct things go wrong, and they are measured by different studies with different methods. Conflating them is why buyers cannot get a straight answer about which tool to trust.

The three failure modes, ranked by how hard they are to catch

The first is a fabricated reference, where the cited source does not exist at all. This is the easiest failure to detect automatically, because a URL either resolves or it does not.

The second is a misattributed claim, where the source exists, resolves, and simply does not say what the report says it says. Detecting this requires reading the source, which is why almost nobody does it.

The third is a real source used out of scope, where a figure is genuine but the sample, the time window or the population has been quietly changed in transit. This is the failure mode that survives every automated check and every skim, and in my experience it is the one that ends up in a board deck.

No major vendor publishes a citation accuracy figure for its deep research product against any of the three. They publish benchmark scores on reasoning and retrieval tasks, which measure whether the agent found the right material, not whether the report you receive represents it correctly. Those are different questions and the gap between them is the entire subject of this post.

The citation accuracy table nobody reads before buying

The most useful public measurement comes from DeepResearch Bench, an arXiv preprint that evaluates agents on 100 tasks written by PhD-level experts across 22 domains. It is not peer reviewed, and that limitation should travel with every number below.

Its FACT framework extracts every statement in a generated report, pairs each with the URL cited beside it, and asks a judge model whether the retrieved page actually supports the statement. Citation accuracy is the proportion of statement and URL pairs that survive that test. Effective citations count how many supported pairs a report contains in absolute terms.

Citation accuracy and effective citations, five deep research agents, 100 tasks
AgentCitation accuracyAvg effective citationsWhat that combination means
Claude-3.7-Sonnet with search93.68%32.48Most reliable per claim. Moderate volume.
Perplexity Deep Research90.24%31.26Strong recall of what it retrieved. Shortest reports.
Grok Deeper Search83.59%8.15Very little supported material per report.
Gemini-2.5-Pro Deep Research81.44%111.21By far the most citations. Roughly 1 in 5 unsupported.
OpenAI Deep Research77.96%40.79Lowest accuracy of the five, mid volume.

Source: DeepResearch Bench, arXiv 2506.11763. Model versions are those tested by the authors and several have since been superseded. Treat the ordering as directional and the specific decimals as belonging to those versions only.

Precision on one axis, volume on the other. No agent wins both. Citation accuracy against average effective citations per report, 100 tasks Citation accuracy 75% 85% 95% 0 60 120 Effective citations Gemini 2.5 Pro DR 81.44% accurate, 111.21 citations OpenAI DR 77.96%, 40.79 Claude w/search 93.68%, 32.48 Perplexity DR 90.24%, 31.26 Grok Deeper Search 83.59%, 8.15
The top left corner is the dangerous quadrant: a long report, densely cited, with the weakest support rate per claim. That is the shape a reader trusts most and should trust least.

Volume is the metric that sells, and it is the wrong one

Gemini 2.5 Pro Deep Research returns 111.21 effective citations per report, more than three times Claude and Perplexity. A 40 page report with 111 sources looks authoritative in a way a 12 page report with 32 sources does not, and buyers respond to that. The accuracy column says one statement in five is not supported by the source sitting next to it.

My position is that volume is actively harmful at this accuracy level. A report with 32 citations at 93.68% contains roughly 2 unsupported claims. A report with 111 citations at 81.44% contains roughly 21. Verification time scales with the number of errors, not with the accuracy rate.

One caution on sample size before anyone quotes that table in a procurement doc. One hundred tasks split across 22 domains leaves fewer than five tasks per domain. That is enough to separate 93% from 78%, and it is not enough to tell you how any of these agents will behave on your specific subject matter. Treat the table as a prior to test, not a verdict to cite.

A second preprint, published in April 2026, takes the narrowest possible definition of a hallucinated reference and applies it at scale. It counts a URL as hallucinated only when the link does not resolve and no Wayback Machine snapshot exists for it at any point in time. A dead link to a page that once existed does not count.

That definition is deliberately conservative, which makes the resulting numbers a floor rather than an estimate. The study covers 53,090 URLs across 10 commercial systems.

Hallucinated and non-resolving URL rates, 10 systems, 53,090 URLs
SystemHallucinated URLsNon-resolving URLs
claude-3-5-sonnet with search3.0%7.8%
claude-3-7-sonnet with search3.2%8.5%
openai-deepresearch3.5%10.1%
gemini-2.5-flash with search4.6%5.4%
gemini-2.5-pro with search4.8%5.9%
gpt-4.15.4%5.4%
gpt-4.1-mini7.4%7.4%
gpt-4o-mini-search-preview8.7%8.7%
gpt-4o-search-preview8.8%8.8%
gemini-2.5-pro-deepresearch13.3%18.5%

Source: arXiv 2604.03173, April 2026. Non-resolving includes links that fail today but have an archive snapshot, so it is the number a reader clicking through will actually experience.

Links that resolve nowhere and were never archived Share of all cited URLs, 53,090 URLs across 10 commercial systems, April 2026 claude-3-5-sonnet w/search3.0% claude-3-7-sonnet w/search3.2% openai-deepresearch3.5% gemini-2.5-flash w/search4.6% gemini-2.5-pro w/search4.8% gpt-4.15.4% gpt-4.1-mini7.4% gpt-4o-mini-search-preview8.7% gpt-4o-search-preview8.8% gemini-2.5-pro-deepresearch13.3% A URL counts as hallucinated only if it fails to resolve and no archive snapshot has ever existed.
Notice which product sits at the bottom. The same vendor's general model with search attached scores 4.8%. The deep research configuration on top of it scores 13.3%.

The same paper runs a second study on 168,021 URLs across 32 academic fields using three current models. There the non-resolving rates are 4.20% for gemini-2.5-pro, 8.47% for gpt-5.1 and 9.38% for claude-sonnet-4-5. The ordering reverses against the first study, which is the clearest possible evidence that these rankings do not survive a model version change.

GhostCite, a February 2026 preprint, asked 13 models to produce citations across 40 computer science domains without retrieval. It then verified 375,440 generated citations against authoritative bibliographic sources. The whole exercise cost roughly $800 in API spend, which is worth noting: this is a study any team could repeat.

The spread is enormous. DeepSeek hallucinated 14.23% of its citations. Hunyuan hallucinated 94.93%. In between sit Claude 4 at 21.84%, GPT-5 at 50.92%, Gemini at 59.47% and Grok 4 at 79.98%.

Retrieval is the intervention. Model choice is the second-order effect. Citation hallucination rate, unassisted generation against search-grounded generation 21.84% 3.2% Claude family 50.92% 3.5% GPT family 59.47% 4.8% Gemini family Unassisted generation, GhostCite Search-grounded, arXiv 2604.03173 Different studies, different model versions and different definitions. Read the gap, not the exact values.
These two sets of bars come from separate papers with separate methods, so the comparison is illustrative rather than measured. The order of magnitude is the finding worth keeping.

The practical consequence lands on anyone wiring a model into an internal research workflow. The retrieval layer is doing almost all of the work on citation quality, and model selection matters only at the margin. Whether the system can cite a document it never opened matters an order of magnitude more, which is the same architectural argument made in the piece on building an internal eval suite.

The contamination has already reached published work

The GhostCite authors also examined 2.2 million citations across 56,381 published papers from 2020 to 2025. They found invalid citations in 1.07% of papers, and an 80.9% increase in 2025 against the 2020 to 2024 average. The base rate is low and the direction is not ambiguous.

Why the rankings invert depending on what you measure

Line the three studies up and no tool wins twice. Claude leads on citation accuracy and on hallucinated URLs in the first study, then places last on non-resolving URLs in the second. Gemini leads on citation volume, places worst on dead links in its deep research configuration, and places best on dead links in its general configuration.

This is not measurement noise. The studies are measuring genuinely different properties, and a product can be built to optimise one at the cost of another. An agent tuned to retrieve aggressively will surface more obscure pages, and obscure pages are more likely to rot.

Confidence does not track accuracy

The Tow Center study is the sharpest evidence on this. Across 1,600 queries the eight tools tested were wrong more than 60% of the time, and ChatGPT Search used hedging language in only 15 of its 134 incorrect answers. Perplexity was the strongest at 37% incorrect. Grok 3 was wrong 94% of the time, and 154 of its 200 citations led to error pages.

The finding I would put in front of a procurement committee is the one about paid tiers. The premium versions answered more prompts correctly and also produced higher error rates, because they declined to answer less often. Paying more bought confidence, not correctness.

There is a shelf-life problem underneath all of this too. OpenAI shut down its dedicated o3-deep-research and o4-mini-deep-research models on 23 July 2026, with gpt-5.6-sol named as the replacement. Every published measurement of "OpenAI Deep Research" now refers to a product you cannot buy, so any comparison table you read, including this one, is a historical document about specific model versions.

What an unverified citation costs downstream

The failure is cheap to make and expensive to catch, which is the worst possible cost structure for a control.

Expert review does not catch it

GPTZero scanned 300 ICLR 2026 submissions in December 2025 and found 50 papers containing at least one confirmed hallucinated citation. Each of those submissions had already been read by three to five peer experts, and most of them missed it. This is vendor-run research from a company selling detection tools, so treat the count as directional rather than settled, but the mechanism it describes is not controversial.

If three domain experts reading a 9 page paper miss a fabricated reference, your manager skimming a 40 page agent-generated report will miss it too. Human review at full volume is not the control here. Sampling is.

Automated verification is also imperfect

The obvious response is to automate the check. CiteCheck, a preprint that grounds each citation against external scholarly sources, reports 88.9% accuracy on a 982-citation benchmark. That is a real improvement over reading everything by hand, and it still means one call in nine is wrong. Any pipeline that treats an automated verifier as ground truth inherits that error rate silently, which is the monitoring problem covered in the piece on detecting agent drift in production.

Where this comparison is weakest

Three things, and the first is the largest.

This post is not original research

The intended version of this article audited every citation in ten reports generated in a single session and published the raw counts. That study was not run. What you are reading instead synthesises four published audits, which means it inherits their sample sizes, their model versions and their definitions rather than controlling any of them. A first-party audit is the right asset here and this is not it, which is the same gap argued in the case for original research as the durable content moat.

The second weakness is that every underlying source is a preprint. DeepResearch Bench, the reference hallucination study, GhostCite and CiteCheck are all arXiv papers that have not completed peer review. At least two of them use a judge model to grade the output of other models. The Tow Center study is the only manually graded source in this post, and it tests AI search tools rather than deep research agents specifically.

The third is that the definitions do not line up. Citation accuracy in DeepResearch Bench asks whether a page supports a statement. Hallucination in the URL study asks whether a page exists. Hallucination in GhostCite asks whether a bibliographic record matches above a similarity threshold. These are three different questions and I have presented them next to each other, which is useful for shape and unsafe for arithmetic. Do not average them.

The audit protocol, and it takes about an hour

The only number that governs your decision is the one measured on your own prompts and in your own domain. Measure it on the model version you are actually paying for. Here is the protocol I would run before signing anything.

Generate three reports per candidate tool on questions you already know the answer to. Sample 10 citations from each report using a random number, not by picking the interesting ones. That gives 30 citations per tool, which is enough to separate a 90% tool from a 78% tool and not enough to separate 90% from 93%.

Four checks per sampled citation, in the order that fails fastest
CheckQuestionTime per citationFailure it catches
1. ResolveDoes the URL load a real page today?5 secondsFabricated and rotted references
2. LocateDoes the specific figure or claim appear on that page?60 secondsMisattribution and drift between adjacent sources
3. ScopeDo the sample, period and population match how the report used it?90 secondsReal numbers used out of context
4. TierIs this the originating source, or a restatement of one?45 secondsSix blog posts tracing to one unpublished survey

Check 3 is the one people skip and the one that matters

Resolution and location are mechanical. Scope requires knowing the subject, which is why it cannot be delegated to a junior reviewer or to an automated verifier. Every failure I have watched reach a decision document passed checks 1, 2 and 4 cleanly.

Record the failure count per tool as a fraction rather than a percentage, because 4 out of 30 is honest about the sample in a way that 13.3% is not. Then set a policy that scales with the stakes: reports feeding an internal discussion get a 10% spot check, and reports feeding a customer commitment or a regulatory filing get 100% verification. The same tiering logic applies to who signs off, which is the argument in the piece on designing human-in-the-loop checkpoints.

Re-run the audit at every model version change. Given that OpenAI retired an entire product line inside 90 days of announcing it, quarterly is the longest defensible interval. The vendor terms governing those changes are compared in the review of model provider contract terms.

Frequently asked questions

Which deep research tool hallucinates the least?

It depends on the failure you care about. On citation accuracy, Claude with search leads at 93.68% and Perplexity Deep Research follows at 90.24%, measured over 100 tasks in DeepResearch Bench. On fabricated URLs, Claude with search and OpenAI Deep Research both sit near 3%, while Gemini 2.5 Pro Deep Research reaches 13.3%. No tool leads on both measures.

How accurate are AI deep research reports?

Published citation accuracy across the five major agents ranges from 77.96% to 93.68%, so between 6% and 22% of cited statements are not supported by the source beside them. Separately, 3% to 13.3% of the URLs those agents produce do not resolve and were never archived. Both figures come from preprints that have not been peer reviewed.

Do AI research agents make up sources?

Yes, though less often than unassisted models do. Without retrieval, tested models fabricated between 14.23% and 94.93% of citations across 375,440 generated references. With search attached, fabricated URL rates fall to roughly 3% to 13%. Retrieval does not eliminate the problem, because a real URL can still be attached to a claim it does not support.

Is Perplexity more accurate than ChatGPT for research?

On the available evidence, yes, on citation fidelity specifically. Perplexity Deep Research scored 90.24% citation accuracy against 77.96% for OpenAI Deep Research in DeepResearch Bench. In the Tow Center test of 1,600 news queries, Perplexity was incorrect 37% of the time against 67% for ChatGPT Search. Perplexity also returns fewer citations per report.

How do I verify AI generated citations?

Sample 10 citations at random per report and run four checks on each. Confirm the URL resolves and confirm the specific claim appears on that page. Then check that the sample and time period match how the report used the figure, and that the page is the originating source rather than a restatement. Budget about 3 minutes per citation and roughly an hour per tool.

Are AI citation accuracy benchmarks reliable?

Treat them as directional. The main studies are arXiv preprints, several use a judge model to grade citations rather than human review, and sample sizes run to 100 tasks in the most cited benchmark. Rankings also invert between studies and reverse on model version changes. They are useful for narrowing a shortlist and insufficient for a final decision.

Where to start

Pick the report your team relied on most heavily in the last month. Sample 10 of its citations at random and run the four checks above. You will have a real number for your own domain in under 40 minutes, and it will be more decision-relevant than every benchmark in this post.

Then write that fraction into your tool evaluation doc next to the price. Right now most teams have the price and nothing at all on the other side of the trade.

Related analysis

Citation behaviour is also becoming a distribution question. See how AI assistants choose which sources to cite in the analysis of citation and rank overlap collapse, and what verified outcomes look like in the piece on case studies that survive checking.

References

  1. Du, Xu, Zhu, Wang and Mao, DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents, arXiv 2506.11763. Used for all citation accuracy and effective citation figures. Preprint.
  2. Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, arXiv 2604.03173, April 2026. Used for all hallucinated and non-resolving URL rates. Preprint.
  3. GhostCite: A Large-Scale Analysis of Citation Validity in the Age of Large Language Models, arXiv 2602.06718, February 2026. Used for unassisted generation rates and the published literature analysis. Preprint.
  4. Jazwinska and Chandrasekar, AI Search Has a Citation Problem, Tow Center for Digital Journalism, Columbia Journalism Review, 6 March 2025. Used for the 1,600 query results and the premium tier finding.
  5. CiteCheck: Retrieval-Grounded Detection of LLM Citation Hallucinations in Scientific Text, arXiv 2605.27700. Used for automated verification accuracy. Preprint.
  6. GPTZero, GPTZero uncovers 50+ hallucinations in ICLR 2026, 5 December 2025. Used for the peer review finding. Vendor research.
  7. OpenAI, API deprecations. Used for the 23 July 2026 shutdown of o3-deep-research and o4-mini-deep-research.

The weakest thing about this source base: five of the seven references are arXiv preprints that have not completed peer review. Two of them use a judge model rather than human graders to score citations. The only manually graded source tests AI search tools, not deep research agents. Every figure belongs to a specific model version, several of which have since been retired.

AV
Aryan Vatsa
Contributing Analyst, Zan Digital. Founding product designer, writing here on how AI products are priced, packaged and measured against each other.

Related reading