From Aryan Vatsa | Product & Market Analysis

Perplexity Enterprise vs Google AI Mode: Only One Is Governable

On this page

Perplexity was the most accurate AI search tool in the largest published citation audit, and it still got 37% of source attributions wrong. Google AI Mode was not in that test, because it did not exist yet. For a research team choosing between Perplexity Enterprise and Google AI Mode, answer quality is not the deciding variable. Source control is.

Key takeaways

  • The most accurate tool anyone has audited still failed 37% of the time. The Tow Center ran 1,600 queries across eight AI search tools in March 2025. More than 60% of all responses were wrong, and Perplexity's 37% was the best score in the group.
  • A working link is not a verified source. A May 2026 preprint testing 14 models found link validity above 94% and topical relevance above 80%, but factual accuracy of only 39% to 77%.
  • Deep research modes fabricate more citations, not fewer. University of Pennsylvania researchers checked 53,090 URLs and found 10.7% hallucinated by deep research agents against 4.8% by search-augmented models.
  • Only one of the two is an administered product. Perplexity Enterprise has an admin console, connector permissions and retention settings. AI Mode is a feature of consumer Search, and the only enterprise control is a Chrome policy that switches it off.
37%Perplexity's incorrect-answer rate, the lowest of eight AI search tools tested. Source: Tow Center, March 2025.
39-77%Factual accuracy of citations across 14 models, against link validity above 94%. Source: arXiv preprint, May 2026.
10.7%Share of deep research agent citation URLs that were hallucinated, against 4.8% for search-augmented models. Source: University of Pennsylvania, April 2026.

The short answer, and who it is for

Perplexity Enterprise is an administered product with an admin console, audit logs and retention controls. Google AI Mode is a feature of consumer Search with none of those. Both retrieve live and both cite. If your team has to defend where a number came from, buy the one an administrator can configure.

This is written for the head of research or head of delivery who has to make a team produce consistently and stand behind the output. The cost difference here is trivial next to the governance difference.

The rest of this post is about why the usual comparison is the wrong one. Almost every published matchup scores these tools on the fluency and usefulness of the answer.

Citation quality: what has actually been measured

Three pieces of independent research carry this section. None of them was funded by a vendor and all three are open to read.

The largest published audit still sets the floor

Klaudia Jaźwińska and Aisvarya Chandrasekar at the Tow Center for Digital Journalism published their audit on 6 March 2025. They took 20 news publishers, picked 10 articles from each, and fed an exact excerpt to eight chatbots with a request for the title, publisher, date and URL. That is 1,600 queries against a known ground truth.

More than 60% of responses contained inaccurate information. Perplexity was wrong 37% of the time and was the best performer. Grok 3 was wrong 94% of the time and produced fabricated or broken URLs in 154 of its 200 responses. Gemini managed a single completely correct answer across the set.

The behaviour underneath those numbers matters more than the ranking. ChatGPT signalled uncertainty just 15 times in 200 responses while misidentifying 134 articles. These systems do not degrade gracefully. They fail at full confidence, which is exactly the failure mode a research workflow is worst at catching.

The best score in the field was 37% wrong. Share of responses containing inaccurate information, across 1,600 queries and eight tools. Perplexity 37% All eight tools over 60% Grok 3 94% 0% 100% Google AI Mode is absent because it had not launched when the audit ran. The Gemini tested here was the standalone app, a different system from AI Mode.
Read the middle bar first. The category average, not the winner, is what your review process has to survive.

A link that resolves is not a source that supports

The Tow Center measured whether the citation pointed at the right article. A May 2026 preprint by Onweller and colleagues asked the harder question: does the cited page actually say the thing. The paper is a preprint and has not been peer reviewed, and that limitation belongs in the same sentence as its findings.

The team parsed inline citations out of generated Markdown reports, then went and retrieved each cited page. They scored three separate properties: does the link load, is the page topically relevant, and does it factually support the claim. Across 14 models, link validity stayed above 94% and relevance above 80%. Factual accuracy came in at 39% to 77%.

That gap is the whole product problem. The blue superscript is there, it clicks through, the page is about the right subject. The part that fails is the part nobody checks, which is whether the sentence you are about to quote is in the document you are about to cite.

One further finding should change how you buy. Factual accuracy fell by roughly 42% on average as tool calls scaled from 2 to 150. More retrieval made attribution worse. I read that as a direct warning about the deep research and agentic modes both vendors are pushing hardest.

Three checks on the same citation, three very different answers. Across 14 models tested in May 2026. Each bar is a stricter test than the one before it. 94%+ Link loads. 80%+ Page is relevant. 39-77% Page supports the claim. 39% Spread across the 14 models. Source: Onweller et al., arXiv preprint 2605.06635, 7 May 2026. Not peer reviewed.
The first two bars are what a reader sees. The third is what an editor finds. Buy for the third one.

Deep research modes fabricate more, not fewer

Delip Rao, Eric Wong and Chris Callison-Burch at the University of Pennsylvania published a large URL audit on 3 April 2026. They extracted every URL from model outputs and classified each one as live, non-resolving, hallucinated, or stale. Hallucinated meant it did not resolve and had no Wayback Machine archive either, which is a careful way to separate invention from ordinary link rot.

The sample is the reason to trust it: 53,090 URLs across ten models on one benchmark, and 168,021 URLs across three models on another covering 32 academic fields. Between 3% and 13% of citation URLs were hallucinated, with 5% to 18% non-resolving overall. Deep research agents pooled at 10.7% hallucinated against 4.8% for search-augmented models.

The mitigation is more interesting than the diagnosis. Adding a deterministic URL health check inside the agent loop cut non-resolving URLs by between 6 and 79 times, to under 1%. That tells you this is a product decision rather than a model limitation.

Freshness is the smaller gap. Corpus reach is the bigger one.

Both products retrieve at query time, so neither is working from a stale training cut-off. AI Mode decomposes a question into subqueries, fires them in parallel against Google's index and live data, and reasons over the passages it gets back. Perplexity runs its own retrieval per query. On recency alone there is no meaningful gap for most research questions.

The real difference is what each system is allowed to see, and both corpora are now developing holes.

The two corpora are diverging, in opposite directions

Perplexity's reach is contested. Cloudflare published a technical investigation on 4 August 2025 documenting an undeclared crawler running 3 to 6 million daily requests alongside the declared crawler's 20 to 25 million. Cloudflare said the undeclared traffic used a generic Chrome user agent and rotated source networks when blocked, and it removed Perplexity from its verified bot list. The Tow Center had already noticed something adjacent: Perplexity's free tier correctly identified all ten excerpts from paywalled National Geographic articles despite the publisher blocking its crawler.

Publishers have moved to court rather than to robots.txt. Encyclopaedia Britannica filed against Perplexity on 15 September 2025, arguing that the answers Perplexity gives are often Britannica's answers. Whatever the legal outcome, the operational effect on a research team is the same. Parts of the web may stop being reachable through this vendor.

Google's corpus was effectively total for a different reason: blocking Googlebot meant leaving Search. That changed on 3 June 2026, when Google gave site owners a Search Console control that removes a site from AI Overviews, AI Mode and AI Overviews in Discover. Google states the control "will not be used as a ranking signal for search results outside of these generative AI Search features" and that opted-out sites "will not receive traffic or impressions from our generative AI features". It began rolling out in the UK. What that trade looks like from the publisher's side is covered in the evidence on GEO versus SEO citation.

So for the first time both corpora are incomplete, and neither vendor publishes a list of what is missing. That is a real limitation on any competitive intelligence workflow, and it is invisible at the interface. The related question of which sources these systems favour when they do have a choice is examined in the analysis of how AI citations track search rank.

Source control is the widest gap

This is where the two products stop being comparable, and it is the section a security review should read first.

What an administrator can pin in Perplexity

Perplexity Enterprise replaces the consumer Focus selector with a source scope. Inside a Space, a user chooses Web, Org Files, Web plus Org Files, or None. That is the control that stops a researcher answering an internal question from the open web, and it is set per workspace rather than per query, so it survives the researcher forgetting.

Internal Knowledge Search puts an uploaded shared repository alongside live web results. Administrators separately govern which app connectors are enabled for the whole organisation, including Google Drive, OneDrive and SharePoint, and can restrict public sharing and external invitations. Retention is configurable rather than fixed, with custom retention and forced deletion available above a seat threshold.

These claims come from Perplexity's own Help Center, which is the primary source and also refused automated retrieval during research for this post.

What a user can prefer in AI Mode

Google's source features are real improvements and they are aimed at a different person. On 6 May 2026 Google moved links inside AI responses so they sit next to the relevant text, added hover previews showing the site name before you click, and began labelling links from publications a user subscribes to. Later in the month, Preferred Sources arrived in AI Overviews and AI Mode.

Every one of those is a per-user preference in a consumer product. A preference is not a control. There is no administrator who can require that a research team has Preferred Sources configured, no way to see whether they did, and no record afterwards of which corpus an answer came from.

Source control, compared at the level a reviewer cares about
CapabilityPerplexity EnterpriseGoogle AI Mode
Restrict answers to approved sourcesScope set per Space: Web, Org Files, both, or None.Preferred Sources, set by each user for themselves.
Search internal documents alongside the webInternal Knowledge Search, over a shared repository.Not available. Search has no tenancy for your files.
Who sets the policyAn administrator, for the whole organisation.The individual, on their own account.
Evidence the policy was appliedAdmin settings plus audit logging.None available to the organisation.

The right-hand column is not a criticism of AI Mode. It is a consumer search feature and it was never sold as a governed research tool. The mistake is putting it in a shortlist against something that was.

The procurement question that decides this

Perplexity Enterprise Pro carries the controls a security review expects: single sign-on, SCIM provisioning, role-based administration, SOC 2 Type II, configurable retention and audit logging. None of that is remarkable. It is the baseline for a tool a company buys.

The surface you cannot administer

AI Mode has no organisation tenancy at all. The Google Workspace admin console has no switch for AI results inside Google Search, because Search is not a Workspace service and your Workspace agreement does not cover it. The control that does exist is a browser policy: Chrome's AIModeSettings takes 0 to leave the feature available and 1 to remove it, deferring to GenAiDefaultSettings when unset.

Being able to turn something off is not the same as being able to run it safely. That policy binds managed Chrome installations, so a researcher on a personal laptop or a phone is outside it entirely. The practical position for most companies is that AI Mode is already in the workflow, ungoverned, and the buying decision is whether to give the team a governed alternative or keep pretending the ungoverned one is not being used. The cost of that gap is quantified in the breakdown of what shadow AI costs when it goes wrong.

What an administrator can actually control. Rows are the controls a research team is asked about in a security review. Perplexity Ent. Google AI Mode. Gemini Ent. Organisation admin console Yes No Yes. Audit log of queries Yes No Yes. Configurable data retention Yes No Yes. Answers scoped to approved sources Yes Per user Yes. Covered by an enterprise agreement Yes No Yes. The middle column is not a worse product, it is a different product class. Verify the right-hand column against current Google Cloud documentation.
Four red cells in the middle column is the finding. AI Mode is not competing with Perplexity Enterprise, it is competing with doing nothing.

What Google actually sells for this job

Google does have an answer for governed research, and it is not AI Mode. The Gemini Enterprise Agent Platform offers grounding with Google Search over public web data, Web Grounding for Enterprise, and grounding against your own data stores, sitting inside Google Cloud identity, logging and VPC Service Controls.

That reframes the shortlist. The honest comparison is Perplexity Enterprise against Gemini Enterprise, two products sold to administrators with contracts behind them. AI Mode belongs in a different row of the analysis, as the thing your team uses when nobody has bought them anything. Treating it as a candidate makes the evaluation look rigorous while comparing objects of different kinds.

If your shortlist has grown to include the general assistant tier from each lab, the admin console differences across those are catalogued in the comparison of ChatGPT, Claude and Gemini enterprise governance. The contract language that should sit under any of them is in the AI contract clauses a CFO should insist on.

Three products, three different things being sold
ProductWhat it isWho the buyer is
Perplexity Enterprise ProA seat-licensed research tool with internal knowledge search and admin controls.A department head with a budget.
Google AI ModeA feature of consumer Google Search, free, with per-user preferences only.Nobody. It arrives by default.
Gemini Enterprise Agent PlatformCloud grounding and agent services under Google Cloud governance.A platform or data team.

Perplexity's Help Center lists Enterprise Pro at 40 dollars per seat per month and 400 per seat per year, with a higher Enterprise Max tier above it. Google publishes no seat price for AI Mode because AI Mode is not sold by the seat.

Where this comparison is weakest

Four things you should hold against everything above.

Nobody has published a head to head test of these two

Every study cited here misses at least one of the products. The Tow Center audit tested Perplexity but predates AI Mode. The Pennsylvania URL audit tested Gemini and OpenAI deep research agents, not Perplexity. The May 2026 attribution paper covers 14 models rather than the two shipping products a buyer actually chooses between. So the citation evidence here supports claims about the category, not a ranking of these two.

The brief for this piece asked for both tools run against identical research briefs. I did not run that test and I will not publish a bake-off I did not conduct. Zan Digital has no first-party evaluation data. The protocol below is the substitute, and it is a worse one than real numbers.

Second, Perplexity's own documentation returned errors to automated retrieval throughout this research, so its feature and price claims here rest on official Help Center pages as surfaced by search rather than pages I read directly. Google's announcements opened without trouble. That asymmetry means the Perplexity column is thinner than the Google column, in a post that concludes in Perplexity's favour.

Third, feature parity in this category moves in weeks. Google shipped three separate link and source changes between May and June 2026 alone. Verify every row against current vendor documentation before you rely on it. Fourth, the strongest case against my conclusion is cost: AI Mode is free and already deployed, and for a team doing low-stakes background reading that may be the correct answer.

Run the identical-brief test yourself

You can settle this for your own team in about a week, and the result will be worth more than any published comparison because it uses your questions.

Write 10 research briefs of the kind your team actually gets. Make five of them questions you already know the answer to and hold the primary document for. Those five are the control, and without them you are grading fluency.

Then score each answer on the three levels the May 2026 paper separated. Does the link load. Is the page about the right subject. Does the page contain the specific claim it was cited for. Score the third one by opening the document, every time, because that is the check nobody does and the one where the tools diverge.

Suggested scoring weights for an identical-brief evaluation
CriterionWeightWhat a passing result looks like
Cited page contains the specific claim40%Verified by opening the source, on every scored claim.
Correct attribution of publisher and date20%Matches the primary document, not a restatement of it.
Source scoping works as configured20%An internal-only query returns nothing from the open web.
Coverage of sources you know exist10%Finds the documents your control questions rest on.
Answer quality and readability10%Clears a bar you set before running the test.

These weights are my judgement, not a measured outcome. They put answer quality last on purpose, because quality converges between these tools and traceability does not. If your output is internal briefing notes rather than published work, raise the last row.

Frequently asked questions

Is Perplexity Enterprise better than Google AI Mode for research?

For a team that has to defend its sources, yes, and the reason is administrative rather than intellectual. Perplexity Enterprise gives an administrator single sign-on, retention settings, connector permissions and the ability to scope answers to approved sources. Google AI Mode is a feature of consumer Search with no organisation tenancy, so nothing a researcher does inside it is configurable or auditable by you.

How accurate are AI search citations?

Worse than the interface suggests. The Tow Center tested 1,600 queries across eight AI search tools in March 2025 and found more than 60 percent of responses wrong. Perplexity was the most accurate at 37 percent incorrect. A May 2026 preprint testing 14 models found link validity above 94 percent but factual accuracy of only 39 to 77 percent.

Can Google Workspace admins turn off AI Mode in Search?

Not from the Workspace admin console, because Google Search is not a Workspace service. The control that exists is a browser policy. Chrome's AIModeSettings policy takes value 0 to leave the feature available and value 1 to remove it, and when unset it defers to GenAiDefaultSettings. That binds managed Chrome installations only, so personal browsers and phones stay outside it.

Can you restrict which sources Perplexity Enterprise uses?

Yes, at both the user and the organisation level. Inside a Space, Enterprise users choose between Web, Org Files, Web plus Org Files, or None, which is how you stop a researcher answering an internal question from the open web. Administrators separately control which connectors are enabled for the organisation, who may share externally, and how long data is retained.

How much does Perplexity Enterprise Pro cost per user?

Perplexity's own Help Center lists Enterprise Pro at 40 dollars per seat per month, or 400 dollars per seat per year, with a discounted rate for eligible educational institutions and nonprofits. A higher Enterprise Max tier exists for organisations running heavier agentic and research workloads. Google publishes no equivalent rate card for AI Mode, because AI Mode is not sold as a seat.

Do deep research modes produce better citations?

The published evidence points the other way. University of Pennsylvania researchers checked 53,090 citation URLs across ten models in April 2026 and found deep research agents hallucinated 10.7 percent of URLs against 4.8 percent for search augmented models. A separate May 2026 preprint found factual accuracy fell about 42 percent as tool calls scaled from 2 to 150.

Where to start this week

Skip the shortlist for a moment. Two things will tell you more than a vendor call.

Pull last month's research output and pick 20 cited claims at random. Open every source and mark whether the document contains the claim it was cited for. If your hit rate is below the 39% floor in the May 2026 paper, your problem is the review step, not the tool, and buying a licence will not fix it.

Then ask both vendors one written question: does your system verify that a cited URL resolves before it puts the citation in front of a user. The Pennsylvania researchers cut non-resolving URLs to under 1% with exactly that check, so it is cheap, and the answer tells you whether the vendor treats attribution as an engineering problem. The failure patterns that follow when nobody asks are collected in the breakdown of how agent pilots actually fail.

Related on buying AI at work

If governance is the constraint rather than the tool, start with the enterprise admin control comparison. If you are drafting the paper, read the contract clauses worth insisting on first.

References

  1. Columbia Journalism Review, Tow Center for Digital Journalism, AI Search Has a Citation Problem, 6 March 2025. Used for the 1,600 query methodology, the 37% and 94% figures and the National Geographic finding.
  2. Onweller, Lumer, Huber, Ramchandani, Subbiah and Feld, Cited but Not Verified: Parsing and Evaluating Source Attribution in LLM Deep Research Agents, arXiv preprint, 7 May 2026. Used for link validity, relevance and factual accuracy across 14 models. Not peer reviewed.
  3. Rao, Wong and Callison-Burch, University of Pennsylvania, Detecting and Correcting Reference Hallucinations in Commercial LLMs and Deep Research Agents, 3 April 2026. Used for the URL sample sizes, hallucination rates and the mitigation result.
  4. Cloudflare, Perplexity is using stealth, undeclared crawlers to evade website no-crawl directives, 4 August 2025. Used for the request volumes, the user agent finding and the de-listing.
  5. Google, How AI Mode and AI Overviews help you explore the web, 6 May 2026. Used for inline links, hover previews and subscription labelling.
  6. Google, New opportunities, control and insights for website owners, 3 June 2026. Used for the Search Console opt-out and the two quoted statements.
  7. Perplexity Help Center, Perplexity Enterprise, accessed August 2026. Used for source scoping, internal knowledge search, connector permissions, retention and pricing.
  8. Chrome Enterprise, AIModeSettings policy, accessed August 2026. Used for the policy values and the fallback to GenAiDefaultSettings.

The weakest part of this source base is asymmetry. Google's announcements and both academic preprints opened directly. Perplexity's Help Center and the Chrome Enterprise policy page returned errors or unrendered content to automated retrieval, so those two rows rest on the official pages as surfaced by search rather than pages read end to end. Two of the three studies are preprints and neither is peer reviewed. No study cited here tested Perplexity Enterprise and Google AI Mode against each other.

AV
Aryan Vatsa
Contributing Analyst, Zan Digital. Founding product designer, writing here on how AI products are priced, packaged and measured against each other.

Related reading