From Sanskriti Khandelwal | Product & Market Analysis
Reddit Is 46.7% of Perplexity's Top Citations. The Playbook Nobody Writes Down
On this page
Reddit accounts for 46.7% of the citations inside Perplexity's ten most-cited sources, the highest single-domain concentration published for any AI engine. That figure has turned Reddit marketing into the most whispered-about tactic in B2B growth. It is also the tactic with the shortest distance between what works and what carries a civil penalty, and most teams cannot tell you where that line sits.
Key takeaways
- The headline number is real, and narrower than it sounds. Reddit is 46.7% of Perplexity's top ten cited sources, not 46.7% of every citation Perplexity makes. Profound measured that across 680 million citations from August 2024 to June 2025.
- The share moves too fast to build a plan on. Semrush tracked ChatGPT citing Reddit in close to 60% of prompt responses in early August 2025, then around 10% by mid September, across more than 230,000 prompts.
- Undisclosed seeding is a legal question, not a taste question. The FTC rule on consumer reviews and testimonials, final since August 2024, reaches reviews written by employees and agents that hide the relationship.
- Getting caught is the base case. A games agency published its own Reddit seeding case study in November 2025 and pulled it after Reddit found it. Moderators of r/Biohackers closed an entire topic to new posts over the same behaviour.
Why Reddit sits at the top of AI citations
The short answer
Reddit is cited heavily because AI engines answer subjective questions, and it holds more first-person product experience than any indexed source. Perplexity leans on it hardest, at 46.7% of its top ten sources. That makes Reddit a real distribution channel, and it makes covert seeding a legal exposure rather than a growth hack.
Two forces put Reddit at the top, and they are worth separating because only one of them is stable.
The supply side: engines reach for forums when the question is subjective
Most commercial queries are not factual lookups. A buyer asking which applicant tracking system is worth the money wants a judgement, and judgements live in threads rather than in documentation. An analysis of 30 million sources by Peec AI, published in March 2026, put Reddit first across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, ahead of YouTube, LinkedIn and Wikipedia.
The same study found Perplexity leaning on Reddit, LinkedIn and G2 for business queries specifically. That combination tells you what the engine is optimising for. It wants named humans describing what happened when they used the thing.
This is the same structural shift covered in the evidence on what separates generative engine optimisation from classic SEO. Retrieval rewards a different asset than ranking did.
The demand side: Reddit is becoming a search engine of its own
Reddit is not passively supplying answers to other people's products. It is building the same surface itself. On the Q2 2026 earnings call, management described Reddit Answers growing to 6 million weekly users from 1 million in December, with more than 70 million weekly users touching Reddit search.
That matters for anyone planning a channel strategy. The threads you take part in now have two retrieval paths, one inside Reddit and one through every engine that cites it. Both paths reward the same thing, which is a comment a real person found useful.
What 46.7% actually measures, and what it does not
This is where most writing about Reddit citations stops being accurate, so it is worth being pedantic.
Top-source share is not total share
Profound's 46.7% is the share Reddit holds within Perplexity's top ten most-cited domains. It is not the share of all Perplexity citations. The same dataset puts Reddit at 21.0% of Google AI Overviews' top sources and 11.3% of ChatGPT's, where Wikipedia takes 47.9% instead.
The distinction changes the strategy. A domain can dominate the leaderboard while the long tail of citations still goes elsewhere. If your category is well covered by trade press and vendor documentation, Reddit may never be the source that answers your buyer's question.
| Engine | Reddit share of top sources | What leads instead | What it implies |
|---|---|---|---|
| Perplexity | 46.7% | Reddit itself, then LinkedIn and G2 on business queries | Community threads are the primary evidence base |
| Google AI Overviews | 21.0% | A wider mix, with Wikipedia at 5.7% | Reddit matters, but it competes with ranked pages |
| ChatGPT | 11.3% | Wikipedia at 47.9% of top sources | Reference material outranks lived experience |
Source: Profound, 680 million citations, August 2024 to June 2025. These are top-source shares, not shares of total citations, and the two numbers behave very differently. Treat the ordering as reliable and the exact percentages as dated.
The number moves week to week
Semrush tracked the top 25 cited domains across more than 230,000 prompts over 13 weeks, from 14 July to 12 October 2025. In that window ChatGPT's Reddit citation rate fell from close to 60% of prompt responses to around 10%. Google AI Mode stayed steady over the same period. Perplexity dipped slightly.
Nobody outside OpenAI knows what caused that. It could be a retrieval change, a licensing change, or a response to spam. What it demonstrates is that a single platform decision can remove most of a channel in six weeks, which is a bad property for a channel you are spending against.
What the rules and the law actually say
Teams tend to assume this is a grey area governed by community etiquette. Part of it is. Part of it is a federal trade regulation with a penalty attached.
The FTC rule reaches insiders and agents
The Federal Trade Commission's final rule on the use of consumer reviews and testimonials, announced on 14 August 2024, bans six categories of conduct. Three of them describe what a Reddit seeding programme does.
Fake reviews and testimonials are prohibited, and the rule names AI-generated reviews and reviews from anyone with no real experience of the product. Insider reviews are prohibited where an officer, manager, employee or agent writes a review without clearly and conspicuously disclosing the relationship to the business. Incentivised reviews are prohibited where compensation is conditioned on a particular sentiment.
Read the word "agent" carefully. Hiring an outside firm to post on your behalf does not move the exposure to the firm. It creates an agent relationship, and the disclosure obligation follows the relationship. That is a clause worth writing into the contract before the campaign, in the same spirit as the clauses finance teams are adding to AI vendor agreements.
Reddit's rules target behaviour, not wording
Reddit's enforcement does not look at whether a comment sounds promotional. It looks at coordination signals: account age, posting patterns across subreddits, voting behaviour and the relationships between accounts. A Reddit spokesperson told 404 Media the company runs systems that detect and prevent inauthentic behaviour, coordinated manipulation and astroturfing.
A polished comment from a three-week-old account with no history in the community is a weaker position than a blunt disclosure from an account with two years of unrelated activity. Reddit is measuring the account, not the copy.
Individual subreddits then apply stricter rules on top, enforced by volunteer moderators and automated tooling. Those rules vary enough that one campaign template applied across ten communities will violate several of them by construction.
Why covert seeding backfires on this platform specifically
Every channel has cheaters. Reddit is unusual in that the audience is also the enforcement mechanism, and it keeps records.
The agency that published its own evidence
In November 2025 a games marketing firm called Trap Plan put a case study on its own site. It described seeding around 100 posts and comments across gaming subreddits to promote a title. An earlier version cited more than 40 posts across communities including r/gaming and r/pcmasterrace. Users found it, archived it, and the pages came down. PC Gamer covered the removal, the chief executive apologised publicly, and the publisher distanced itself from the work.
The mechanics of that failure are the part to study. The campaign was not detected by Reddit's systems. It was detected because the agency needed a case study to sell the next campaign, and a case study is discoverable. Any vendor good enough to be worth hiring has the same incentive.
When a community notices, it closes. The more expensive failure is not the individual ban. It is the community deciding your entire category is spam. Moderators of r/Biohackers placed a moratorium on new posts about peptides and hormone replacement therapy after companies in that space flooded the subreddit, as reported by 404 Media in June 2026.
Consider what that does to a legitimate operator in the same category. The threads that would have cited them no longer exist. The subreddit engines were pulling from has stopped producing material. One competitor's seeding programme removed the channel for everyone, and no amount of budget reopens it.
The research explains the mechanism, and it is uncomfortable. Cornell researchers Hal Triedman, Tingwei Zhang and Vitaly Shmatikov tested whether user-generated content can steer deep research agents. Working in a sandbox using Reddit API content rather than live posts, they found a snippet as short as 13 words was enough to push agent output toward spam or scam content. They also reported that deep research agents cite user-generated content in roughly half of queries, and that close to 25% of all citations came from those sites.
This is a preprint, tested in simulation, and it has not been through peer review. Both limits belong in the same sentence as the finding. The direction is still the point: manipulation is cheap, which is why platforms and regulators will keep tightening, and why a strategy built on it has a short life.
The participation playbook that survives contact
Here is what I would actually do, stated as a position rather than a menu. Treat Reddit as a support channel that happens to be indexed, not as a publishing channel that happens to have users.
Answer questions, do not plant posts
The unit of value on Reddit is a specific answer to a specific question from someone qualified to give it. That is also the unit AI engines extract. A founder answering a pricing question with real numbers, from an account carrying the company name in the bio, is doing the compliant version and the effective version at the same time.
Disclosure belongs in the account, not in a footnote at the end of a long comment. Put the employer in the profile and name it in the first line when the question touches your product. The FTC standard is clear and conspicuous, and a profile a reader has to click is arguably neither.
Publish the data, then let other people carry it
The strongest Reddit position is not being present in the thread. It is being the thing the thread cites. Original research with a stated sample, window and exclusion list gets posted by other people, and those posts are the ones engines pick up. That advantage compounds in a way commenting does not, which is the argument in the piece on which data actually functions as a moat.
This route is slower, and it is the only version that does not depend on nobody looking too closely. It also produces an asset you own, rather than a comment the platform can remove.
Then buy the ad, and run the AMA. The unglamorous answer is that Reddit sells advertising, and it is the overwhelming majority of what the company earns. Q2 2026 advertising revenue was $762 million against $43 million of data licensing. Paid placement, official AMAs and a brand profile are the sanctioned routes, and they carry no risk of a moderator naming you in a pinned post.
| Tactic | Verdict | Reason |
|---|---|---|
| Named employee answering questions, employer disclosed in profile and comment | Works | Compliant with the FTC insider provision and survives moderator scrutiny |
| Publishing original data others choose to post | Works, slowly | Produces the citation without requiring you to be in the thread |
| Reddit ads, official AMAs, brand profiles | Works, and it is boring | Sanctioned surfaces, and the company's actual business model |
| Asking customers to post, with an unconditional thank-you | Conditional | Legal only where the incentive is not tied to sentiment or to posting at all |
| Employees recommending the product without disclosure | Prohibited | Squarely inside the FTC undisclosed insider review category |
| Agency-run seeded posts from unaffiliated-looking accounts | Prohibited and traceable | The agent relationship carries the disclosure duty, and case studies leak |
| Honest assessment of the grey version | It does work, short term | This is the concession. Seeding produces citations, which is why the market exists. It also transfers the durable risk to you and the fee to the agency. |
What to measure, and what to ignore
Most Reddit reporting measures the wrong object. Upvotes and impressions tell you about the post. They tell you nothing about whether an engine repeated you.
Three measurements are worth the effort. First, run your ten highest-intent buying queries through Perplexity, ChatGPT and Google AI Mode monthly, and record which domains appear rather than only whether you do. Second, track whether the Reddit threads that mention you are ones you took part in or ones you did not, because the second number is the health metric. Third, record the date of every reading, because these systems change underneath you.
Do not report citation counts as a headline without the query set attached. A citation rate is meaningless without the questions that produced it. The overlap between AI citations and classic rankings is also thinner than most dashboards suggest, as covered in the analysis of where citation and rank stop agreeing. The discipline is the same one applied to any AI spend in the piece on where measurable return has actually appeared.
What getting caught actually costs
The risk is usually described as reputational, which understates it. There are four separate enforcement mechanisms and they do not overlap.
| Who enforces | What triggers it | What it costs you |
|---|---|---|
| The FTC | Reviews or endorsements from employees or agents without clear disclosure | Civil penalties, plus orders to compensate consumers for harm |
| Coordinated accounts, inauthentic personas, vote manipulation | Shadowbans and account removal, taking every thread you built with them | |
| Subreddit moderators | Any promotional pattern the community notices | Pinned warnings naming the brand, and topic-wide moratoriums |
| The public archive | A deleted case study, a screenshot, a moderator post | Trade press coverage attached to the brand name indefinitely |
The fourth row is the one that gets ignored in planning and dominates in practice. Reddit threads and news coverage about a brand being caught are themselves highly cited sources, so the penalty enters the same retrieval systems the campaign was aimed at.
That last point deserves stating plainly. If the reason you want Reddit citations is that engines trust Reddit, then you also have to accept that engines will trust the Reddit thread accusing you of astroturfing. The channel does not discriminate in your favour.
Where this argument is weakest
Three genuine problems with everything above.
The measurement is vendor research, not independent audit. Profound, Semrush and Peec AI all sell tools that benefit from marketers believing AI citations matter. Their numbers are the best available and they are not disinterested. None of the three publishes a methodology detailed enough for an outside party to reproduce, and the query sets that produced these shares are not public.
Treat the ordering as reliable, since three independent datasets agree that Reddit leads. Treat any specific percentage as directional and dated. The honest version of the headline is that Reddit is the most-cited domain by a wide margin, not that it is exactly 46.7%.
Reddit itself says the guarantee does not exist. Reddit pushed back on the vendor pitch, telling 404 Media there is no guarantee that what is on Reddit reaches large language models, and questioning the legitimacy of any provider claiming otherwise. That is a self-interested statement from a company with a data licensing business to protect. It is also correct on the mechanics, because no vendor controls the retrieval layer.
Anyone selling a Reddit citation as a deliverable is selling an outcome they cannot produce. The relationship between presence and citation remains a correlation nobody has isolated, a problem examined in the comparison of ChatGPT citations against Google rankings.
The commercial case for the grey version is stronger than I want it to be. Seeding works in the short run. It is cheaper than original research, faster than building a two-year account history, and enforcement is inconsistent. A founder with 18 months of runway is making a defensible risk calculation when they choose it, and calling that choice obviously stupid would be dishonest.
My disagreement is about the asymmetry rather than the ethics. The upside is a citation share that a single retrieval change can remove, as ChatGPT's 2025 drop showed. The downside is a permanent public record. Those are not symmetric bets, and that is the argument, not the moralising.
Frequently asked questions
Why does Perplexity cite Reddit so much?
Perplexity is tuned to answer subjective and comparative questions, and Reddit holds more first-person product experience than any other indexed source. Profound's analysis of 680 million citations found Reddit made up 46.7% of Perplexity's ten most-cited sources, against 21.0% on Google AI Overviews and 11.3% on ChatGPT. On ChatGPT, Wikipedia leads instead at 47.9% of top sources.
Is Reddit marketing against the rules?
Participating openly is not. Reddit's rules target inauthentic behaviour, coordinated accounts and vote manipulation rather than promotional wording. What crosses the line is concealment. Posting from accounts that hide an employment or agency relationship breaches Reddit's user agreement, and in the United States it also falls under the FTC rule on undisclosed insider reviews, which carries civil penalties.
Can you pay an agency to get your brand mentioned on Reddit?
You can, and the exposure stays with you. The FTC rule covers reviews written by agents of a business, so hiring an outside firm creates the disclosure obligation rather than transferring it. There is also a practical problem. Agencies publish case studies to win the next client, and those case studies are how the most-reported seeding campaign of the past year was discovered.
Does posting on Reddit actually get you cited by ChatGPT?
Sometimes, and nobody can promise it. Reddit told 404 Media there is no guarantee that content on Reddit reaches large language models, and questioned any provider claiming to deliver that. Citation shares also move sharply. Semrush measured ChatGPT citing Reddit in close to 60% of responses in early August 2025 and around 10% by mid September, across more than 230,000 prompts.
What is astroturfing on Reddit and how do brands get caught?
Astroturfing means running accounts that appear to be independent users while promoting a product. Brands get caught three ways: Reddit's systems flag coordination patterns across accounts, moderators recognise repeated messaging and pin warnings, and agencies publish case studies that users find and archive. In June 2026, r/Biohackers closed an entire topic to new posts because of promotional flooding.
How should a B2B brand use Reddit in 2026?
Treat it as a support channel that happens to be indexed. Have named employees answer specific questions from profiles that disclose the employer, publish original data with a stated sample so other people post it for you, and use Reddit's paid surfaces where you need reach. Measure whether threads you did not take part in mention you, because that number is the real health signal.
Where to start this week
Two things, and the first one takes an hour.
Run your ten highest-intent buying queries through Perplexity and ChatGPT, and write down every domain that appears alongside whether Reddit is among them. If Reddit is absent from eight of ten, this whole topic is a distraction for your category and you have just saved a quarter of effort.
Then audit what you already have running. Ask any agency touching community channels for the account names they post from, in writing. If they will not give you the list, you have found the exposure, and the contract clause you need is a disclosure requirement rather than a performance target.
Related on AI search
This sits alongside the evidence on what actually differs between generative engine optimisation and SEO and the measurement of how far citation and ranking have drifted apart.
References
- Profound, AI platform citation patterns, analysis of 680 million citations, August 2024 to June 2025. Used for all top-source share figures.
- Search Engine Land, AI search engines cite Reddit, YouTube and LinkedIn most, 31 March 2026, reporting Peec AI's analysis of 30 million sources. Used for the cross-engine ranking.
- Semrush, The most-cited domains in AI: a three-month study, 230,000+ prompts, 14 July to 12 October 2025. Used for the ChatGPT volatility figures.
- Federal Trade Commission, Final rule banning fake reviews and testimonials, 14 August 2024. Used for all regulatory claims, including the insider and agent provisions.
- 404 Media, It is trivially easy to use Reddit to manipulate AI search, research suggests, 15 June 2026. Used for the Cornell preprint findings and the Reddit spokesperson statement.
- PC Gamer, Game marketing company takes down blog post bragging about how good it is at astroturfing Reddit, November 2025. Used for the Trap Plan case.
- AdExchanger, AI SEO has brands astroturfing Reddit, June 2026, summarising 404 Media reporting. Used for the r/Biohackers moratorium and Reddit's response to vendor claims. This citation should be upgraded to the 404 Media original.
- CNBC, Reddit Q2 2026 earnings report, 30 July 2026. Used for revenue, advertising and data licensing figures.
The weakest thing about this source base: three of the four citation datasets come from companies selling AI visibility tooling, and none publishes a query set an outside party could reproduce. The regulatory and financial figures are primary. The citation percentages are not, and they are dated to the windows stated above.
Related reading