From Aryan Vatsa | Product & Market Analysis
GEO vs SEO: 45 Studies Later, the Citation Checklist Barely Holds Up
On this page
Most GEO advice traces back to one 2024 paper that tested content rewrites inside a fixed context window. When later researchers rebuilt that test as a full retrieval pipeline, the same rewrites cut final citations by about 6%. GEO vs SEO is not two disciplines. It is one pipeline with three stages, and almost every checklist sold today optimises the wrong one.
Key takeaways
- The famous 40% GEO uplift was measured with retrieval switched off. The original benchmark added statistics and quotations to documents already sitting in the model's context, then counted word share in the answer. Nothing in that design tests whether you get retrieved.
- Peer review has not replicated it. C-SEO Bench tested 9 methods across 6 domains at NeurIPS 2025. Only 3 of 54 method and domain pairs produced a statistically significant gain, and none worked for question answering.
- Position in the context window beat every content trick tested. Moving a document to first position scored 2.77 in the retail task against 0.36 for the best content-side method. That is a ranking problem, which is what SEO has always been.
- The largest measured citation signal sits off your own site. Across 75,000 brands, branded web mentions correlated with AI Overview presence at 0.664. Backlinks managed 0.218.
What GEO and SEO actually optimise
Generative engine optimisation means editing content so an AI answer uses it. Answer engine optimisation is the same practice under a different label. Search engine optimisation means getting a page indexed, retrieved and ranked.
Framed that way they sound like rivals. They are not. An AI answer comes out of a pipeline with three stages, and the two disciplines act on different stages.
Stage one is retrieval. The engine issues a query and pulls a candidate set of documents. Stage two is reranking, which decides the order those documents occupy inside the model's context. Stage three is generation, where the model writes prose and attaches citations.
SEO acts on stages one and two. Almost every GEO tactic acts on stage three. That distinction carries the whole argument, because a stage-three tactic pays nothing if stages one and two never happened.
The blunt version: you cannot be quoted from a document the engine never fetched. Every measured GEO gain in the published literature is conditional on the document already being in the context window. This is the part the category name hides, and it is why the shift toward generated answers is a distribution question before it is a formatting question, as the piece on software after the dashboard argues in a different setting.
The 40% figure comes from one synthetic benchmark
The number underneath most GEO advice comes from a single paper. Aggarwal and colleagues published GEO: Generative Engine Optimization at KDD 2024. They built GEO-Bench, 10,000 queries drawn from nine datasets, and tested 9 content modifications against a baseline.
The result quoted everywhere is "up to 40%". The paper's own table is more specific and more useful. Position-adjusted word count rose from a 19.3 baseline to 27.2 for quotation addition and 25.2 for statistics addition. Cite sources reached 24.6 and fluency optimisation 24.7.
Keyword stuffing scored 17.7, below the baseline. That row is the one worth keeping. The methods that helped were the ones that made a document easier to extract from. The method that gamed the string did worse than doing nothing at all.
One further finding is almost never quoted. Adding citations lifted visibility for the fifth-ranked site by 115.1%, while the top-ranked site lost 30.3% on average. GEO redistributes attention inside a fixed set of documents. It does not expand the set.
What the paper tested, and what it did not
The evaluation used GPT-3.5-turbo as the generator and the cleaned text of the top 5 Google results as the context. The document being optimised was already in that context by construction, on every single query.
So the experiment answers one question well. Given that your page has been retrieved, do these edits make the model quote more of it? It answers nothing about whether the page gets retrieved in the first place.
The authors said so. Their stated limitations include that the black-box design prevented any evaluation of effects on search rankings. The industry kept the 40% and dropped that sentence.
Replication has not been kind to the checklist
A benchmark result is a hypothesis. Two years of follow-up work has been testing this one, and the results have moved against the original.
3 of 54, and none for question answering
C-SEO Bench, published at NeurIPS Datasets and Benchmarks in 2025, tested 9 conversational SEO methods across 6 domains and two tasks. Out of 54 method and domain combinations, 3 produced a statistically significant gain.
The authors' summary is blunt. Most current methods are, in their words, not only largely ineffective but frequently harmful to document ranking. For the question answering task they found no method effective at all.
The comparison that matters sits in the same paper. Moving a document to first position in the model's context scored 2.77 in the retail domain against 0.36 for the best content-side method. Ordinary ranking work outperformed the best GEO tactic by roughly 7 times.
They also modelled adoption directly. Gains fall steadily as more competitors adopt the same method, approaching zero at full adoption. That is a congested, zero-sum game, and it has the same shape as every other tactic available to everyone, which is the argument running through the piece on what actually makes a wrapper defensible.
Body-only rewrites can cost you the retrieval you already had
The July 2026 critical survey of this field reviewed 45 studies published between November 2023 and July 2026. It reports a pipeline experiment in which body-only optimisation reduced top-20 presence by about 9%, top-10 presence after reranking by about 16%, and final citation by about 6%.
Read that against the 41% headline. Same class of edit, opposite sign, once retrieval is switched back on. A rewrite tuned for quotability can make a page less like the query it was meant to answer.
The survey's own conclusion is the sentence I would pin above a content team's desk. Within the reviewed corpus, no technique showed a stable, longitudinal, cross-platform causal effect on organic discoverability. Conditional effects on already-retrieved content are well established. Organic effects are not.
The signal that holds up sits off your own website
If content edits are this weak, something else must explain why certain brands appear constantly in AI answers. The largest correlational dataset published so far points off-site.
Ahrefs analysed 75,000 brands in May 2025 and measured Spearman correlations against presence in Google AI Overviews. Branded web mentions correlated at 0.664. Branded anchors reached 0.527 and branded search volume 0.392. Backlinks, the currency of 20 years of SEO practice, came in at 0.218, below Domain Rating at 0.326.
The author states plainly that correlation is not causation, and that caveat is load-bearing here. Large brands get mentioned more, and they also do everything else more. No published dataset separates those two things.
It is still the largest single effect anyone has measured, and the mechanism it suggests is coherent. A model that has seen your company discussed by independent sources has an entity to attach a recommendation to. A model that has only seen your own marketing has a document.
My reading is that the durable asset here is not a content format at all. It is being the source other people cite, which is close to what proprietary data buys you in the argument about data moats against workflow moats. It is also why this looks less like a search update and more like distribution moving to a new gatekeeper, the pattern described in the piece on where agent marketplaces capture value.
The checklist that survives the evidence
Five practices clear the bar. The bar is a measured effect from a source that published its method and its sample.
| Practice | Best available evidence | What it does not show |
|---|---|---|
| Rank, and stay retrievable | C-SEO Bench: first position in context scored 2.77 against 0.36 for the best content method | One benchmark, two tasks. Does not transfer automatically to every engine. |
| Put the answer in the first block | Citation absorption study: high-influence pages carried 12.50 times more headings and 8.94 times higher list density than the bottom quartile | Observational. Cannot separate structure from length, topic or quality. |
| Carry extractable evidence: statistics, quotations, definitions | KDD 2024: 19.3 to 27.2 for quotations. Absorption study: definitions showed 57.33% higher influence | Both effects are conditional on the page already being retrieved. |
| Earn third-party mentions | Ahrefs, 75,000 brands: 0.664 correlation with AI Overview presence | Correlational, single vendor, single engine. No causal test exists. |
| Measure with repetition | Daily source-level overlap of 0.34 to 0.42, and 57.8% of repeated ChatGPT prompts did not trigger a web search | Not a growth tactic. It is the floor beneath any claim you make. |
Rows 2, 3 and 5 come from preprints published in 2026. Row 1 is peer reviewed. Row 4 is vendor data with the vendor's own causation caveat attached.
Notice what that list is. Four of the five are things a competent search team was already doing in 2019. The genuinely new item is the measurement discipline, and it exists only because the engines are unstable enough to fool you.
The instability is worth sitting with. If more than half of repeated prompts never trigger a web search, the model is answering from memory, and no on-page edit reaches it. Engines also have a cost reason to search less often, which is the pressure described in the piece on what inference costs do to AI margins.
There is a sixth item I cannot yet evidence and would do anyway. Publish original data. The absorption study found encyclopedia-style pages averaged 0.2144 influence against 0.0726 for news media, and original figures are the one input a competitor cannot copy off your page. That is also the only version of this work with a measurable business case, in the sense argued in the breakdown of where AI return has actually shown up.
Three things to stop paying for
Each of these is sold as GEO. Each has been measured, and each measurement came back empty.
llms.txt
Ahrefs checked 137,210 domains with traffic in May 2026. 28% published a valid llms.txt file. Of those, 97% received zero requests during the month.
The sharper finding is the negative case. On domains where the file did not exist, requests to that path were 98% human and zero AI bot. AI tools do not go looking for a file that is not there, so publishing one does not create demand for it.
Google's John Mueller has said since June 2025 that no AI system uses it. The idea is reasonable and the adoption is absent. Cheap to add, and it returns nothing you can measure.
Schema markup sold as a citation lever
Google's own documentation on AI features and your website is unusually direct. A page must be indexed and eligible to be shown with a snippet, and there are no additional technical requirements.
It goes further. You do not need to create new machine readable files, AI text files, or markup to appear in these features, and there is no special schema.org structured data to add.
Schema still earns rich result eligibility in classic search, which is worth having on its own terms. Buying it as an AI citation lever is buying a different product from the one it is.
Question-shaped headings on every section
The instinct is to convert every body heading into the question a buyer types. The citation absorption study measured the opposite of the expected effect: pages in question and answer format showed 5.74% lower influence than non-Q&A pages.
One preprint is not a verdict. It is enough to stop treating question headings as free upside. Keep query phrasing in the FAQ and the opening answer block, where the retrieval case for it is strongest.
| Sold as | What the measurement found | Verdict |
|---|---|---|
| llms.txt for AI crawlers | 97% of published files got zero requests in May 2026; 404s drew zero AI bot traffic | No measurable return. Skip it. |
| Schema markup for AI citation | Google states no special structured data is needed for AI features | Keep for rich results. Do not buy it as a citation lever. |
| Question headings throughout the body | Q&A format pages showed 5.74% lower influence in a 2026 study | Confine query phrasing to the FAQ and opening block. |
Where this argument is weakest
This post argues that GEO is mostly SEO with a narrow extraction layer bolted on. Here is the strongest case against that.
The best evidence for GEO is better than my summary implies
Watanabe and Nakayashiki at Glasp ran a log-based natural experiment on a real domain. They applied an optimisation bundle to one corpus in January 2026 and used the untreated remainder of the same domain as a contemporaneous control.
ChatGPT referrals to treated pages grew 6.1 times against 3.5 times for the control. Their interrupted time series estimated a level increase of 1.82 times above the platform trend, with a 95% confidence interval of 1.31 to 2.54.
That is first-party, measured on a live site, with a control group. It is the strongest single result in favour of doing this work, and I have given it three paragraphs against four sections of scepticism. Weigh it accordingly.
The authors are more careful than most vendors. Their placebo-in-time test returned p=0.16, so they call the effect suggestive rather than conclusive. Their treated pages were already trending upward before the change, and the four tactics were bundled, so nobody knows which one worked.
Citation is not traffic, and the flywheel claim did not hold
The same study found no support for the stronger claim that AI citations lift organic search rankings. Organic clicks to the treated pages fell about 25% over the period, in line with a site-wide fall of about 20%.
So the honest position is narrow. A well-executed overhaul can plausibly close to double your AI referral share above the platform trend. It will not rescue your search traffic, and most of the growth you see in AI referrals is the platform growing rather than you.
The last weakness is the evidence base itself. Most of what I have cited is under 18 months old, most of it is preprint rather than peer reviewed, and the engines change faster than the papers do. Anyone claiming a settled answer here, in either direction, is running ahead of the data.
Frequently asked questions
What is the difference between GEO and SEO?
SEO optimises whether a page is retrieved and ranked. GEO, or generative engine optimisation, optimises how a retrieved page is used inside an AI answer. They are stages of one pipeline rather than rival disciplines. The evidence so far is that stage one still dominates: in C-SEO Bench, moving a document to first position in the model's context scored 2.77 in the retail task against 0.36 for the best content-side method.
Does GEO actually work?
Sometimes, and less reliably than the marketing suggests. The original 2024 GEO paper reported gains of up to 40%, measured on documents already sitting in the model's context. A 2025 benchmark published at NeurIPS tested 9 methods across 6 domains and found only 3 of 54 combinations produced a statistically significant gain. None worked for question answering, and gains shrank as more sites adopted the same method.
Do statistics and quotations really improve AI citations?
They help once your page is already retrieved. In the 2024 GEO benchmark, adding quotations raised position-adjusted word count from 19.3 to 27.2, and adding statistics raised it to 25.2. Keyword stuffing scored 17.7, below the baseline. Later work found the same edits can reduce retrieval when applied to the body alone, cutting top-20 presence by about 9%. Add evidence because it is true, not because it is a trick.
Does schema markup help you get cited by ChatGPT?
There is no published evidence that it does. Google's own documentation states that a page needs only to be indexed and eligible for a snippet, with no additional technical requirements, and says you do not need special schema.org structured data for AI features. Schema still earns rich result eligibility in classic search, which is a real benefit. Treat it as hygiene rather than a citation lever, and do not expand types hoping for a lift.
Is llms.txt worth adding to my site?
On the current evidence, no. Ahrefs checked 137,210 domains with traffic in May 2026 and found 97% of published llms.txt files received zero requests that month. Files that returned a 404 drew no AI bot traffic at all, which means AI tools are not looking for them. Google's John Mueller has said repeatedly that no AI system uses the file. It costs little, and it returns nothing measurable.
How do I measure whether AI search is citing my content?
Run the same prompt many times, not once. Commercial engines are unstable: one 2026 audit recorded daily source-level overlap of 0.34 to 0.42, and found 57.8% of repeated ChatGPT prompts did not trigger a web search at all. 7 to 8 repetitions per prompt is a sensible floor. Record the prompt set, the date, the engine and the account state, because all four move the answer.
Where to start this week
Two moves, in this order.
First, take your 10 highest-value commercial pages and check one thing on each. Does the page answer its own title inside the first 60 words, with a number and a named source? That is the single content-side edit with evidence behind it at both the retrieval stage and the citation stage.
Second, build the measurement floor before you buy any tool. Write down 25 prompts a real buyer would type. Run each 8 times on a logged-out session and record which domains got cited. Repeat it monthly, on the same day, with the same prompts.
Without that baseline you cannot separate a retainer's effect from the platform's own growth, which is precisely the confound the Glasp researchers had to design around. And if a vendor's pitch cannot name which pipeline stage it operates on, you already have your answer.
Related on this site
The distribution question underneath all of this is covered in where agent marketplaces capture value, and the defensibility question in what actually makes a wrapper defensible.
References
- Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization, ACM SIGKDD 2024. Used for all position-adjusted word count figures and the rank redistribution result.
- Puerto, Gubri, Green, Oh and Yun, C-SEO Bench: Does Conversational SEO Work?, NeurIPS Datasets and Benchmarks 2025. Used for the 3 of 54 result and the 2.77 against 0.36 comparison.
- Martinez, Optimizing Visibility in Generative Engines: A Critical Survey of Generative Engine Optimization 2023 to 2026, 15 July 2026. Used for the pipeline-stage losses, the engine instability figures and the synthesis quote.
- Zhang, He and Yao, From Citation Selection to Citation Absorption, 29 April 2026. Used for heading and list density, definitions, Q and A format, and domain-type influence.
- Watanabe and Nakayashiki, Disentangling Answer Engine Optimization from Platform Growth, 11 August 2026. Used for the 1.82 multiplier, the confidence interval and the organic traffic finding.
- Ahrefs, An Analysis of AI Overview Brand Visibility Factors, 26 May 2025. Used for every Spearman correlation figure.
- Ahrefs, We Analyzed 137K Sites: 97% of llms.txt Files Never Get Read, 15 June 2026. Used for all llms.txt adoption and request figures.
- Google Search Central, AI Features and Your Website. Used for the structured data and eligibility statements.
The weakest thing about this source base: only 2 of the 8 items have been through peer review. Three are preprints under 6 months old, two are published by a company that sells search software, and one is written by the platform being examined. Figures are current as of 21 August 2026.
Related reading