From Sanskriti Khandelwal | Product & Market Analysis

Reddit Is 46.7% of Perplexity's Top Citations. The Playbook Nobody Writes Down

On this page

Reddit accounts for 46.7% of the citations inside Perplexity's ten most-cited sources, the highest single-domain concentration published for any AI engine. That figure has turned Reddit marketing into the most whispered-about tactic in B2B growth. It is also the tactic with the shortest distance between what works and what carries a civil penalty, and most teams cannot tell you where that line sits.

Key takeaways

  • The headline number is real, and narrower than it sounds. Reddit is 46.7% of Perplexity's top ten cited sources, not 46.7% of every citation Perplexity makes. Profound measured that across 680 million citations from August 2024 to June 2025.
  • The share moves too fast to build a plan on. Semrush tracked ChatGPT citing Reddit in close to 60% of prompt responses in early August 2025, then around 10% by mid September, across more than 230,000 prompts.
  • Undisclosed seeding is a legal question, not a taste question. The FTC rule on consumer reviews and testimonials, final since August 2024, reaches reviews written by employees and agents that hide the relationship.
  • Getting caught is the base case. A games agency published its own Reddit seeding case study in November 2025 and pulled it after Reddit found it. Moderators of r/Biohackers closed an entire topic to new posts over the same behaviour.
46.7%Reddit's share of Perplexity's top ten cited sources. Source: Profound, 680 million citations, 2024 to 2025.
$43MReddit's quarterly data licensing revenue, about 5% of its $805M total. Source: CNBC on Reddit's Q2 2026 results.
13 wordsSnippet length Cornell researchers found enough to steer a research agent toward spam output. Source: 404 Media, June 2026.

Why Reddit sits at the top of AI citations

The short answer

Reddit is cited heavily because AI engines answer subjective questions, and it holds more first-person product experience than any indexed source. Perplexity leans on it hardest, at 46.7% of its top ten sources. That makes Reddit a real distribution channel, and it makes covert seeding a legal exposure rather than a growth hack.

Two forces put Reddit at the top, and they are worth separating because only one of them is stable.

The supply side: engines reach for forums when the question is subjective

Most commercial queries are not factual lookups. A buyer asking which applicant tracking system is worth the money wants a judgement, and judgements live in threads rather than in documentation. An analysis of 30 million sources by Peec AI, published in March 2026, put Reddit first across ChatGPT, Google AI Mode, Gemini, Perplexity and AI Overviews, ahead of YouTube, LinkedIn and Wikipedia.

The same study found Perplexity leaning on Reddit, LinkedIn and G2 for business queries specifically. That combination tells you what the engine is optimising for. It wants named humans describing what happened when they used the thing.

This is the same structural shift covered in the evidence on what separates generative engine optimisation from classic SEO. Retrieval rewards a different asset than ranking did.

The demand side: Reddit is becoming a search engine of its own

Reddit is not passively supplying answers to other people's products. It is building the same surface itself. On the Q2 2026 earnings call, management described Reddit Answers growing to 6 million weekly users from 1 million in December, with more than 70 million weekly users touching Reddit search.

That matters for anyone planning a channel strategy. The threads you take part in now have two retrieval paths, one inside Reddit and one through every engine that cites it. Both paths reward the same thing, which is a comment a real person found useful.

What 46.7% actually measures, and what it does not

This is where most writing about Reddit citations stops being accurate, so it is worth being pedantic.

Top-source share is not total share

Profound's 46.7% is the share Reddit holds within Perplexity's top ten most-cited domains. It is not the share of all Perplexity citations. The same dataset puts Reddit at 21.0% of Google AI Overviews' top sources and 11.3% of ChatGPT's, where Wikipedia takes 47.9% instead.

The distinction changes the strategy. A domain can dominate the leaderboard while the long tail of citations still goes elsewhere. If your category is well covered by trade press and vendor documentation, Reddit may never be the source that answers your buyer's question.

Reddit's concentration by engine, share of top ten cited sources
EngineReddit share of top sourcesWhat leads insteadWhat it implies
Perplexity46.7%Reddit itself, then LinkedIn and G2 on business queriesCommunity threads are the primary evidence base
Google AI Overviews21.0%A wider mix, with Wikipedia at 5.7%Reddit matters, but it competes with ranked pages
ChatGPT11.3%Wikipedia at 47.9% of top sourcesReference material outranks lived experience

Source: Profound, 680 million citations, August 2024 to June 2025. These are top-source shares, not shares of total citations, and the two numbers behave very differently. Treat the ordering as reliable and the exact percentages as dated.

Reddit's share of each engine's ten most-cited sources Profound, 680 million citations, August 2024 to June 2025 Perplexity46.7% Google AI Overviews21.0% ChatGPT11.3% On ChatGPT, Wikipedia holds 47.9% of top sources. The engine you optimise for decides the asset you build.
Notice the spread, not the top bar. A tactic tuned for Perplexity is close to irrelevant on the engine with the most users.

The number moves week to week

Semrush tracked the top 25 cited domains across more than 230,000 prompts over 13 weeks, from 14 July to 12 October 2025. In that window ChatGPT's Reddit citation rate fell from close to 60% of prompt responses to around 10%. Google AI Mode stayed steady over the same period. Perplexity dipped slightly.

Nobody outside OpenAI knows what caused that. It could be a retrieval change, a licensing change, or a response to spam. What it demonstrates is that a single platform decision can remove most of a channel in six weeks, which is a bad property for a channel you are spending against.

One platform decision removed most of the channel in six weeks Share of ChatGPT prompt responses citing Reddit, two measured points from a 13-week study about 60% Early August 2025 about 10% Mid September 2025 -50 points Source: Semrush, 230,000+ prompts, 14 July to 12 October 2025. Intermediate weeks are not plotted.
Only the two endpoints are measured figures. The dashed line marks the direction of travel, not a weekly series.

What the rules and the law actually say

Teams tend to assume this is a grey area governed by community etiquette. Part of it is. Part of it is a federal trade regulation with a penalty attached.

The FTC rule reaches insiders and agents

The Federal Trade Commission's final rule on the use of consumer reviews and testimonials, announced on 14 August 2024, bans six categories of conduct. Three of them describe what a Reddit seeding programme does.

Fake reviews and testimonials are prohibited, and the rule names AI-generated reviews and reviews from anyone with no real experience of the product. Insider reviews are prohibited where an officer, manager, employee or agent writes a review without clearly and conspicuously disclosing the relationship to the business. Incentivised reviews are prohibited where compensation is conditioned on a particular sentiment.

Read the word "agent" carefully. Hiring an outside firm to post on your behalf does not move the exposure to the firm. It creates an agent relationship, and the disclosure obligation follows the relationship. That is a clause worth writing into the contract before the campaign, in the same spirit as the clauses finance teams are adding to AI vendor agreements.

Reddit's rules target behaviour, not wording

Reddit's enforcement does not look at whether a comment sounds promotional. It looks at coordination signals: account age, posting patterns across subreddits, voting behaviour and the relationships between accounts. A Reddit spokesperson told 404 Media the company runs systems that detect and prevent inauthentic behaviour, coordinated manipulation and astroturfing.

A polished comment from a three-week-old account with no history in the community is a weaker position than a blunt disclosure from an account with two years of unrelated activity. Reddit is measuring the account, not the copy.

Individual subreddits then apply stricter rules on top, enforced by volunteer moderators and automated tooling. Those rules vary enough that one campaign template applied across ten communities will violate several of them by construction.

Why covert seeding backfires on this platform specifically

Every channel has cheaters. Reddit is unusual in that the audience is also the enforcement mechanism, and it keeps records.

The agency that published its own evidence

In November 2025 a games marketing firm called Trap Plan put a case study on its own site. It described seeding around 100 posts and comments across gaming subreddits to promote a title. An earlier version cited more than 40 posts across communities including r/gaming and r/pcmasterrace. Users found it, archived it, and the pages came down. PC Gamer covered the removal, the chief executive apologised publicly, and the publisher distanced itself from the work.

The mechanics of that failure are the part to study. The campaign was not detected by Reddit's systems. It was detected because the agency needed a case study to sell the next campaign, and a case study is discoverable. Any vendor good enough to be worth hiring has the same incentive.

When a community notices, it closes. The more expensive failure is not the individual ban. It is the community deciding your entire category is spam. Moderators of r/Biohackers placed a moratorium on new posts about peptides and hormone replacement therapy after companies in that space flooded the subreddit, as reported by 404 Media in June 2026.

Consider what that does to a legitimate operator in the same category. The threads that would have cited them no longer exist. The subreddit engines were pulling from has stopped producing material. One competitor's seeding programme removed the channel for everyone, and no amount of budget reopens it.

The research explains the mechanism, and it is uncomfortable. Cornell researchers Hal Triedman, Tingwei Zhang and Vitaly Shmatikov tested whether user-generated content can steer deep research agents. Working in a sandbox using Reddit API content rather than live posts, they found a snippet as short as 13 words was enough to push agent output toward spam or scam content. They also reported that deep research agents cite user-generated content in roughly half of queries, and that close to 25% of all citations came from those sites.

This is a preprint, tested in simulation, and it has not been through peer review. Both limits belong in the same sentence as the finding. The direction is still the point: manipulation is cheap, which is why platforms and regulators will keep tightening, and why a strategy built on it has a short life.

The participation playbook that survives contact

Here is what I would actually do, stated as a position rather than a menu. Treat Reddit as a support channel that happens to be indexed, not as a publishing channel that happens to have users.

Answer questions, do not plant posts

The unit of value on Reddit is a specific answer to a specific question from someone qualified to give it. That is also the unit AI engines extract. A founder answering a pricing question with real numbers, from an account carrying the company name in the bio, is doing the compliant version and the effective version at the same time.

Disclosure belongs in the account, not in a footnote at the end of a long comment. Put the employer in the profile and name it in the first line when the question touches your product. The FTC standard is clear and conspicuous, and a profile a reader has to click is arguably neither.

Publish the data, then let other people carry it

The strongest Reddit position is not being present in the thread. It is being the thing the thread cites. Original research with a stated sample, window and exclusion list gets posted by other people, and those posts are the ones engines pick up. That advantage compounds in a way commenting does not, which is the argument in the piece on which data actually functions as a moat.

This route is slower, and it is the only version that does not depend on nobody looking too closely. It also produces an asset you own, rather than a comment the platform can remove.

Then buy the ad, and run the AMA. The unglamorous answer is that Reddit sells advertising, and it is the overwhelming majority of what the company earns. Q2 2026 advertising revenue was $762 million against $43 million of data licensing. Paid placement, official AMAs and a brand profile are the sanctioned routes, and they carry no risk of a moderator naming you in a pinned post.

Where Reddit's money comes from, Q2 2026 Total revenue $805 million. The AI licensing line is the small one. Advertising $762M 94.7% of revenue, growing 64% year on year $43M licensing Data licensing grew 24% year on year and is about 5% of the total. Source: Reddit Q2 2026 results, reported 30 July 2026.
The company monetises attention through ads, not through the AI licensing line everyone talks about. Worth knowing before you assume seeding is the only route in.
Tactics, and where each one sits
TacticVerdictReason
Named employee answering questions, employer disclosed in profile and commentWorksCompliant with the FTC insider provision and survives moderator scrutiny
Publishing original data others choose to postWorks, slowlyProduces the citation without requiring you to be in the thread
Reddit ads, official AMAs, brand profilesWorks, and it is boringSanctioned surfaces, and the company's actual business model
Asking customers to post, with an unconditional thank-youConditionalLegal only where the incentive is not tied to sentiment or to posting at all
Employees recommending the product without disclosureProhibitedSquarely inside the FTC undisclosed insider review category
Agency-run seeded posts from unaffiliated-looking accountsProhibited and traceableThe agent relationship carries the disclosure duty, and case studies leak
Honest assessment of the grey versionIt does work, short termThis is the concession. Seeding produces citations, which is why the market exists. It also transfers the durable risk to you and the fee to the agency.

What to measure, and what to ignore

Most Reddit reporting measures the wrong object. Upvotes and impressions tell you about the post. They tell you nothing about whether an engine repeated you.

Three measurements are worth the effort. First, run your ten highest-intent buying queries through Perplexity, ChatGPT and Google AI Mode monthly, and record which domains appear rather than only whether you do. Second, track whether the Reddit threads that mention you are ones you took part in or ones you did not, because the second number is the health metric. Third, record the date of every reading, because these systems change underneath you.

Do not report citation counts as a headline without the query set attached. A citation rate is meaningless without the questions that produced it. The overlap between AI citations and classic rankings is also thinner than most dashboards suggest, as covered in the analysis of where citation and rank stop agreeing. The discipline is the same one applied to any AI spend in the piece on where measurable return has actually appeared.

What getting caught actually costs

The risk is usually described as reputational, which understates it. There are four separate enforcement mechanisms and they do not overlap.

Four enforcement mechanisms, four different costs
Who enforcesWhat triggers itWhat it costs you
The FTCReviews or endorsements from employees or agents without clear disclosureCivil penalties, plus orders to compensate consumers for harm
RedditCoordinated accounts, inauthentic personas, vote manipulationShadowbans and account removal, taking every thread you built with them
Subreddit moderatorsAny promotional pattern the community noticesPinned warnings naming the brand, and topic-wide moratoriums
The public archiveA deleted case study, a screenshot, a moderator postTrade press coverage attached to the brand name indefinitely

The fourth row is the one that gets ignored in planning and dominates in practice. Reddit threads and news coverage about a brand being caught are themselves highly cited sources, so the penalty enters the same retrieval systems the campaign was aimed at.

That last point deserves stating plainly. If the reason you want Reddit citations is that engines trust Reddit, then you also have to accept that engines will trust the Reddit thread accusing you of astroturfing. The channel does not discriminate in your favour.

Where this argument is weakest

Three genuine problems with everything above.

The measurement is vendor research, not independent audit. Profound, Semrush and Peec AI all sell tools that benefit from marketers believing AI citations matter. Their numbers are the best available and they are not disinterested. None of the three publishes a methodology detailed enough for an outside party to reproduce, and the query sets that produced these shares are not public.

Treat the ordering as reliable, since three independent datasets agree that Reddit leads. Treat any specific percentage as directional and dated. The honest version of the headline is that Reddit is the most-cited domain by a wide margin, not that it is exactly 46.7%.

Reddit itself says the guarantee does not exist. Reddit pushed back on the vendor pitch, telling 404 Media there is no guarantee that what is on Reddit reaches large language models, and questioning the legitimacy of any provider claiming otherwise. That is a self-interested statement from a company with a data licensing business to protect. It is also correct on the mechanics, because no vendor controls the retrieval layer.

Anyone selling a Reddit citation as a deliverable is selling an outcome they cannot produce. The relationship between presence and citation remains a correlation nobody has isolated, a problem examined in the comparison of ChatGPT citations against Google rankings.

The commercial case for the grey version is stronger than I want it to be. Seeding works in the short run. It is cheaper than original research, faster than building a two-year account history, and enforcement is inconsistent. A founder with 18 months of runway is making a defensible risk calculation when they choose it, and calling that choice obviously stupid would be dishonest.

My disagreement is about the asymmetry rather than the ethics. The upside is a citation share that a single retrieval change can remove, as ChatGPT's 2025 drop showed. The downside is a permanent public record. Those are not symmetric bets, and that is the argument, not the moralising.

Frequently asked questions

Why does Perplexity cite Reddit so much?

Perplexity is tuned to answer subjective and comparative questions, and Reddit holds more first-person product experience than any other indexed source. Profound's analysis of 680 million citations found Reddit made up 46.7% of Perplexity's ten most-cited sources, against 21.0% on Google AI Overviews and 11.3% on ChatGPT. On ChatGPT, Wikipedia leads instead at 47.9% of top sources.

Is Reddit marketing against the rules?

Participating openly is not. Reddit's rules target inauthentic behaviour, coordinated accounts and vote manipulation rather than promotional wording. What crosses the line is concealment. Posting from accounts that hide an employment or agency relationship breaches Reddit's user agreement, and in the United States it also falls under the FTC rule on undisclosed insider reviews, which carries civil penalties.

Can you pay an agency to get your brand mentioned on Reddit?

You can, and the exposure stays with you. The FTC rule covers reviews written by agents of a business, so hiring an outside firm creates the disclosure obligation rather than transferring it. There is also a practical problem. Agencies publish case studies to win the next client, and those case studies are how the most-reported seeding campaign of the past year was discovered.

Does posting on Reddit actually get you cited by ChatGPT?

Sometimes, and nobody can promise it. Reddit told 404 Media there is no guarantee that content on Reddit reaches large language models, and questioned any provider claiming to deliver that. Citation shares also move sharply. Semrush measured ChatGPT citing Reddit in close to 60% of responses in early August 2025 and around 10% by mid September, across more than 230,000 prompts.

What is astroturfing on Reddit and how do brands get caught?

Astroturfing means running accounts that appear to be independent users while promoting a product. Brands get caught three ways: Reddit's systems flag coordination patterns across accounts, moderators recognise repeated messaging and pin warnings, and agencies publish case studies that users find and archive. In June 2026, r/Biohackers closed an entire topic to new posts because of promotional flooding.

How should a B2B brand use Reddit in 2026?

Treat it as a support channel that happens to be indexed. Have named employees answer specific questions from profiles that disclose the employer, publish original data with a stated sample so other people post it for you, and use Reddit's paid surfaces where you need reach. Measure whether threads you did not take part in mention you, because that number is the real health signal.

Where to start this week

Two things, and the first one takes an hour.

Run your ten highest-intent buying queries through Perplexity and ChatGPT, and write down every domain that appears alongside whether Reddit is among them. If Reddit is absent from eight of ten, this whole topic is a distraction for your category and you have just saved a quarter of effort.

Then audit what you already have running. Ask any agency touching community channels for the account names they post from, in writing. If they will not give you the list, you have found the exposure, and the contract clause you need is a disclosure requirement rather than a performance target.

Related on AI search

This sits alongside the evidence on what actually differs between generative engine optimisation and SEO and the measurement of how far citation and ranking have drifted apart.

References

  1. Profound, AI platform citation patterns, analysis of 680 million citations, August 2024 to June 2025. Used for all top-source share figures.
  2. Search Engine Land, AI search engines cite Reddit, YouTube and LinkedIn most, 31 March 2026, reporting Peec AI's analysis of 30 million sources. Used for the cross-engine ranking.
  3. Semrush, The most-cited domains in AI: a three-month study, 230,000+ prompts, 14 July to 12 October 2025. Used for the ChatGPT volatility figures.
  4. Federal Trade Commission, Final rule banning fake reviews and testimonials, 14 August 2024. Used for all regulatory claims, including the insider and agent provisions.
  5. 404 Media, It is trivially easy to use Reddit to manipulate AI search, research suggests, 15 June 2026. Used for the Cornell preprint findings and the Reddit spokesperson statement.
  6. PC Gamer, Game marketing company takes down blog post bragging about how good it is at astroturfing Reddit, November 2025. Used for the Trap Plan case.
  7. AdExchanger, AI SEO has brands astroturfing Reddit, June 2026, summarising 404 Media reporting. Used for the r/Biohackers moratorium and Reddit's response to vendor claims. This citation should be upgraded to the 404 Media original.
  8. CNBC, Reddit Q2 2026 earnings report, 30 July 2026. Used for revenue, advertising and data licensing figures.

The weakest thing about this source base: three of the four citation datasets come from companies selling AI visibility tooling, and none publishes a query set an outside party could reproduce. The regulatory and financial figures are primary. The citation percentages are not, and they are dated to the windows stated above.

MJ
Sanskriti Khandelwal
Contributing Analyst, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading