From Ritu Raj | Product & Market Analysis
Nearly 90% of ChatGPT Citations Go to Pages Google Ranks Below 21
On this page
Semrush measured it. When ChatGPT search cites a page, that page ranks 21 or worse in Google organic almost 90% of the time. The easy reading is that Google rank stopped mattering. The data does not support that reading, and the narrower one it does support is more useful.
Key takeaways
- The 90% figure is scored against the wrong index. Seer Interactive found 87% of SearchGPT citations matched Bing's top organic results while only 56% matched Google's. The median Google rank of a cited page was 17.
- Two large datasets contradict each other in public. Semrush puts almost 90% of cited pages at Google position 21 or worse. AirOps, across 548,534 retrieved pages, found 55.8% of cited pages ranked inside the top 20.
- The traits that separate cited pages are on-page and cheap. Title to query overlap above 50% lifted the citation rate from 9.3% to 20.1%. Pages of 500 to 2,000 words beat pages over 5,000.
- Domain authority is the trait that does not work. Close to 74% of AirOps citations went to sites with authority under 80, and the 80 to 100 band had the lowest citation rate of any tier at 15.0%.
What the 90% figure actually measures
Do Google rankings still decide ChatGPT citations? Mostly not, at the query you care about. Semrush found that ChatGPT search cites pages ranking 21 or worse in Google organic almost 90% of the time. Ranking first still helps. It is no longer the gate it was for a decade.
The number comes from Semrush research published on 21 July 2025, covering more than 500 digital marketing topics and subtopics. The same study reported that a visit from AI search was worth 4.4 times the average organic visit on conversion rates.
Read the claim slowly. It says the cited page does not rank in Google's top 20 for the related query. It does not say the page is thin, obscure, or unranked anywhere. Those are three different failures and the study measures none of them.
That distinction carries the rest of this post. A page can sit at position 40 on Google and still be the strongest answer available to a system that is not consulting Google.
The pages were ranked. Just not by Google
ChatGPT search does not read Google's index. For most of its working life it read Bing's, and it now supplements that with OpenAI's own crawl through the OAI-SearchBot agent. Scoring its output against Google is scoring one shop's stock against a different shop's shelves.
The retrieval layer is a separate market with a separate index, in the same way the assistant is a separate surface from the page it summarises. That second shift is the subject of the piece on software after the interface stops being a page.
87% matched Bing, 56% matched Google
Seer Interactive ran the test the other way round in February 2025. It put 100 queries to SearchGPT, tracked the identical queries in both engines, and joined more than 500 citations back to the results by exact URL.
87% of the citations matched Bing's top organic results, most of them inside the top 10. Only 56% matched Google's. The median Google rank of a cited page was 17, and the average was 28.
Those numbers describe the same set of pages. On one index they are page-one results. On the other they sit on page two and beyond. So I would stop calling them pages Google buries. They are pages Google ranks differently, retrieved by a system that never asked Google.
Fan-out is why both numbers can be true
The second reason the 90% figure overstates the break is that ChatGPT rarely searches the query you typed. AirOps analysed 548,534 pages retrieved across 15,000 prompts in March 2026. 89.6% of those prompts triggered two or more follow-up searches, expanding 15,000 prompts into 43,233 queries.
32.9% of cited pages appeared only in results for a fan-out query, and 95% of those queries had no recorded monthly search volume. A page cited off one of them looks unranked to any tool tracking the original keyword, because the keyword it ranked for is not in the tool.
Score the same corpus against every query, original and fan-out together, and the picture inverts. AirOps found 55.8% of cited pages ranked in the top 20 for at least one. Pages at Google position 1 were cited 43.2% of the time, roughly 3.5 times the rate of pages outside the top 20.
Two credible datasets sit about 35 points apart on the same relationship. That is a disagreement about what counts as the query, and it is unsettled. The parallel case inside Google's own answers, where top-10 overlap fell from 76.1% to 37.9% in eight months, is measured in the breakdown of the AI Overview rank overlap collapse.
| Study | Sample and date | Headline finding |
|---|---|---|
| Semrush, AI search traffic | 500 plus topics, July 2025 | ChatGPT cites pages at Google position 21 or worse almost 90% of the time |
| Seer Interactive | 100 queries, 500 plus citations, February 2025 | 87% of citations matched Bing's top results, 56% matched Google's |
| AirOps, retrieval and fan-out | 548,534 pages, 15,000 prompts, March 2026 | 85% of retrieved pages never cited. 55.8% of cited pages ranked in Google's top 20 |
| AirOps, on-page factors | 353,799 pages, 16,851 queries, April 2026 | Heading relevance was the strongest on-page factor. 500 to 2,000 words performed best |
| Semrush, content optimisation | 304,805 cited URLs, 11,882 prompts, January 2026 | 5 of 13 content qualities separated cited pages. Freshness was not one |
Every row is vendor research. These 5 publish their samples and windows, which is why they are here and a dozen widely quoted figures are not. Rows 1 and 3 contradict each other, and that is discussed rather than averaged away.
Trait 1: the heading answers the question in the reader's words
If rank is a weak predictor of citation, something else is doing the work. The measured answer is unglamorous. Cited pages say the answer where a machine can find it, in the words the machine searched for.
Title overlap more than doubles the citation rate
AirOps found pages with 50% or greater overlap between title and query were cited 20.1% of the time. Below that threshold the rate was 9.3%. That is a 2.2 times lift from phrasing alone.
A second AirOps study in April 2026, covering 353,799 pages across 16,851 unique queries run three times each, put heading relevance ahead of every other on-page factor it tested. Pages with a strong heading to query match were cited 41.0% of the time, against roughly 30% for a weak match.
This is the least fashionable finding in the category and the most actionable one. It costs nothing to phrase a section heading as the thing a buyer actually asks.
It carries a trap. The standard this blog writes to bans question-shaped body headings, because a page of them reads as templated. The way out is to keep the heading declarative and put the answer inside it. The heading above this paragraph carries the query terms and the finding, without turning the page into a quiz.
Trait 2: narrow pages beat comprehensive ones
The April 2026 study measured word count against citation rate. Pages between 500 and 2,000 words performed best. Pages over 5,000 words did worse than pages under 500.
Four to 10 subheadings was the optimal structural band. Content aged 30 to 89 days outperformed both fresher pages and pages more than 2 years old.
The mechanism is easy to guess and hard to prove. A retrieval system is hunting for a passage, not a document. A 6,000-word guide contains the answer and dilutes it, and the model has to decide which of 40 sections you meant.
Where this cuts against our own house style
This blog publishes at 2,300 words and up. That sits at the very top of the band the data prefers, and by the definition above every post here is a comprehensive guide.
I am not going to pretend that is costless. If the AirOps finding holds, long analytical posts trade citation rate for depth. The honest position is that we have chosen depth, and that the choice has a measurable price.
The mitigation is structural. Sections that answer one question completely, with the answer in the first sentence, give a retriever a short page inside a long one. Whether that fully substitutes has not been measured by anyone.
Trait 3: structure the model can lift out
Semrush ran the largest of these studies in January 2026. It compared 304,805 URLs cited by ChatGPT Search, Google AI Mode and Perplexity against 921,614 URLs that ranked in Google's top 20 and were not cited, across 11,882 prompts.
Five content qualities separated the two groups. Clarity and summarisation led at 32.83%, then experience and expertise signals at 30.64%, question and answer formatting at 25.45%, section structure at 22.91% and structured data elements at 21.60%.
Read that list as a single instruction. Put a clean answer near the top, say who wrote it and why they would know, and mark the structure up so a parser can find the boundaries.
Structured data earns a caveat. The April study found pages carrying JSON-LD were cited 38.5% of the time against 32.0% without it. That is correlation on observational data, and schema travels with a dozen other quality signals. It is hygiene worth having, not a lever worth a project.
The seven signals that separated nothing
The more useful half of the Semrush result is the negative half. Of 13 parameters tested, 7 showed little to no ability to distinguish cited pages from uncited ones.
Freshness was among them. So were readability, information density and quote usage. Every one of those sits in the standard advice list for AI search visibility, and this dataset says they do not discriminate.
Freshness is the interesting failure, because the April AirOps study found the opposite and reported an optimal age band of 30 to 89 days. One measured association across a cited-versus-uncited split, the other measured citation rate by age. Treat freshness as unsettled, and do not rebuild a publishing calendar on it.
Trait 4: domain authority is not one of the traits
The trait most content budgets are actually buying is authority. It is the one the data does not support.
In the March 2026 AirOps dataset, close to 74% of citations went to sites with a domain authority under 80. The 20 to 40 band contributed 26.0% of citations. The 80 to 100 band contributed 25.4%, a smaller share from a far more established set of sites.
Citation rates say it more sharply. Sites between 0 and 80 were cited at 21.5% to 23.6%. Sites between 80 and 100 were cited at 15.0%, the lowest rate of any tier. The most authoritative domains were the least likely to be cited once retrieved.
I would not read that as authority being penalised. High-authority domains rank for far more queries, so they get pulled into retrieval for many they do not answer well. The conclusion survives the caveat. Buying authority does not buy citations, which is a familiar error: paying for a proxy because the real asset is harder to build, as in the difference between a data moat and a workflow moat.
Third-party surfaces beat your own domain
A synthesis published by 5W Public Relations in May 2026, drawing on nine vendor datasets covering January 2025 to April 2026, put Wikipedia at 13.15% and Reddit at 11.97% of US ChatGPT citations. Together they take more than a quarter of everything cited. The Wall Street Journal, The New York Times, Bloomberg and the Financial Times do not appear in the top 20 at all.
The same synthesis reported roughly a 3 times citation increase for brands present on G2, Capterra, Trustpilot and Yelp. It is a synthesis rather than a study, so treat the multiplier as directional. The direction held across all nine inputs.
The implication for a software company is uncomfortable and cheap to act on. The page most likely to get you cited is one you do not own, cannot edit, and have never budgeted for.
The source mix moves faster than a content plan
Everything above describes a system that can change without notice. It already has.
Semrush tracked more than 230,000 prompts in weekly snapshots for 13 weeks, from 14 July to 12 October 2025, across ChatGPT search, Google AI Mode and Perplexity.
In early August, ChatGPT cited Reddit in close to 60% of prompt responses and Wikipedia in roughly 55%. By mid-September Reddit was near 10% and Wikipedia had fallen below 20%. Both held steady on Google AI Mode and Perplexity across the identical weeks, which tells you the change was one platform's decision.
What a buyer should take from a 4-day collapse
The timing lines up with Google removing its num=100 search parameter on 11 September 2025, which cut off deep-result access for many tools. Semrush's own head of organic and AI visibility offered a different reading, that ChatGPT was deliberately reducing over-citation of a narrow set of domains.
Search Engine Journal covered the drop and concluded the cause is not fully established. Nobody outside OpenAI can settle it, and nobody outside OpenAI was told it was happening.
The lesson is not about Reddit. It is that this distribution channel has one operator, no changelog and no notice period. A content plan whose return depends on a single citation surface carries counterparty risk, which is the same problem as handing distribution to an agent marketplace you do not control.
| Common advice | What the data says | Confidence |
|---|---|---|
| Rank on page one of Google to get cited | Position 1 pages are cited at 43.2%, yet 90% of citations sit at position 21 or worse | Contested. Two datasets disagree by about 35 points |
| Match headings and titles to the query | Title overlap above 50% lifts citation rate from 9.3% to 20.1% | Reasonable. Two studies, one publisher |
| Publish longer, more comprehensive pages | 500 to 2,000 words performed best. Over 5,000 did worse than under 500 | Reasonable, single study |
| Publish more often to stay fresh | Freshness showed no distinguishing effect in the Semrush split | Unsettled. AirOps found the opposite |
| Build domain authority | 74% of citations went to sites under authority 80. The top tier was cited at 15.0% | Reasonable, single study, plausible confound |
Where this argument is weakest
The case above rests entirely on commercial research into a system whose operators publish nothing. That is a serious limitation and it deserves more than a footnote.
Every dataset here is vendor research
Semrush, Seer Interactive, AirOps and 5W all sell services that become more valuable if AI search visibility is a discipline you need help with. None of this work is peer reviewed. None publishes raw data for replication.
That does not make it wrong. It does mean you should trust the sign of an effect more than its magnitude, and read every percentage here as directional.
The methods differ in ways that matter. The April AirOps study scraped the ChatGPT interface rather than calling the API. That is closer to what a real user sees and considerably harder for anyone else to reproduce.
The two headline numbers also still contradict each other. Almost 90% at position 21 or worse cannot describe the same world as 55.8% inside the top 20. I have offered fan-out as the reconciliation, and that remains a hypothesis rather than a finding.
Retrieval is not citation
The quietest number in the AirOps work is the most important one. ChatGPT left 85% of the pages it retrieved uncited.
Every trait in this post was measured on pages that had already made it into the retrieval set. None of it tells you how to get retrieved. Retrieval is where index coverage, crawl access and, yes, ranking still do most of the work.
So the practical reading is narrower than the headline. Rank gets you into the room. What is on the page decides whether you get quoted.
Nothing here measures revenue either. A citation is not a click, and a click is not a trial. Checking what a channel returned before funding more of it is the discipline applied in the piece on where measurable AI return has shown up.
Frequently asked questions
Do Google rankings still matter for ChatGPT citations?
Less than they did, and not in the way most reporting suggests. Semrush found ChatGPT cites pages ranking 21 or worse on Google almost 90% of the time. AirOps found pages ranking first in Google were cited at 58.4%, against 14.2% at position 10. Rank still predicts retrieval. It no longer predicts which retrieved page gets quoted, and those are separate stages.
Why does ChatGPT cite pages that rank low on Google?
Because it is not reading Google. ChatGPT search draws on Bing and on OpenAI's own crawl. Seer Interactive found 87% of SearchGPT citations matched Bing's top organic results against 56% for Google. It also expands your prompt into follow-up searches. AirOps found 32.9% of cited pages appeared only in results for one of those expansions, and 95% of them had no recorded search volume.
What kind of pages does ChatGPT cite most?
Community and reference sources dominate. A 5W Public Relations synthesis of nine vendor datasets put Wikipedia at 13.15% and Reddit at 11.97% of US ChatGPT citations, more than a quarter between them. The Wall Street Journal, The New York Times, Bloomberg and the Financial Times were absent from the top 20. Brand-owned pages compete for what is left of the citation pool.
How long should a page be to get cited by ChatGPT?
AirOps measured 353,799 pages in April 2026 and found the 500 to 2,000 word band performed best. Pages over 5,000 words were cited less often than pages under 500. Four to 10 subheadings was the optimal structure. Treat that as a finding about passages rather than documents: a retrieval system wants one section that answers one question completely.
Does domain authority help you get cited in AI search?
The measured answer is no. In the AirOps March 2026 dataset, close to 74% of citations went to sites with domain authority under 80. Sites in the 80 to 100 band had the lowest citation rate of any tier at 15.0%, against 21.5% to 23.6% below it. High-authority domains get retrieved for many queries they do not answer well, which is the likeliest explanation.
How do I track whether ChatGPT is citing my site?
There is no first-party report. OpenAI publishes no equivalent of Search Console, and referral traffic from assistants is frequently stripped of its referrer. The practical options are server log analysis for OAI-SearchBot and ChatGPT-User hits, and paid visibility trackers that run prompt panels. Both are proxies. Log data tells you about retrieval, prompt panels tell you about a sample of queries you chose.
Where to start this week
Two things, both doable without a new budget line.
First, take your 5 highest-value pages and rewrite the section headings so each states an answer in the words a buyer would use. Keep them declarative. Then leave the pages alone for 60 days and check whether any assistant cites them for that phrasing.
Second, find out where your category is discussed on surfaces you do not own. Review platforms, Reddit threads, the relevant Wikipedia entries. Read them before deciding whether to act, because a clumsy intervention there is worse than none.
What I would not do is commission an authority-building programme on the strength of AI search. The data in this post says that money buys retrieval you may already have, and not the citation you are paying for.
Related analysis
The same break measured inside Google's own answers is in the AI Overview rank overlap collapse, and what actually defends a position is in the piece on the wrapper insult.
References
- Semrush, We Studied the Impact of AI Search on SEO Traffic, 21 July 2025. Used for the 90% figure and the 4.4 times visit value.
- Seer Interactive, 87% of SearchGPT Citations Match Bing's Top Results, 6 February 2025. Used for the Bing and Google match rates and median rank.
- AirOps, The Influence of Retrieval, Fan-out, and Google SERPs on ChatGPT Citations, 12 March 2026. Used for retrieval, fan-out, top-20 overlap and domain authority figures.
- Search Engine Land, ChatGPT citations reward ranking and precision over length, 16 April 2026. Used for heading match, word count, JSON-LD and content age.
- Semrush, How We Built a Content Optimization Tool for AI Search, 14 January 2026. Used for the 13 parameters and the 5 that separated cited pages.
- Semrush, The Most-Cited Domains in AI: A 3-Month Study, 2025. Used for the Reddit and Wikipedia citation collapse and the platform comparison.
- Search Engine Journal, Why Reddit's ChatGPT Citation Drop Isn't Fully Explained, 2025. Used for the contested explanation of the collapse.
- 5W Public Relations, Citation Source Audit Q1 2026, 11 May 2026. Used for the Wikipedia and Reddit shares and the review platform multiplier.
The weakest thing about this source base: all 8 references trace to companies selling AI visibility tooling or communications services, none is peer reviewed, and references 1 and 3 report findings that contradict each other. Reference 8 is a synthesis of nine other vendor datasets, and reference 6 is read from a published chart.
Related reading