From Aryan Vatsa | Product & Market Analysis

Original Research Is the Only Content Moat Left, and a Study Costs About $457

On this page

Everything summarisable you publish is now a commodity, because a machine can restate it in one sentence. A number you collected yourself cannot be restated by anything that did not collect it. That is why original research is the only content moat left. The entry price is about $457 for a 200-response study.

Key takeaways

  • A 200-response study costs about $457 in panel fees. Prolific's published corporate rate is participant pay plus a 42.8% platform fee. At the recommended $12 per hour, an 8-minute survey pays $1.60 per completed response.
  • Credibility is a disclosure problem before it is a budget problem. AAPOR asks for 11 elements on public release, and five of them fit into a single paragraph printed under your chart.
  • The cheapest defensible format is not a survey at all. Counting something in public documents costs nothing but time, and anyone who doubts the count can rerun it against the same filings.
  • The evidence that statistics lift AI citation does not test whose statistics they are. The GEO paper inserted statistics into existing pages and measured the visibility change. Provenance was never a variable in that experiment.
$457Panel cost of a 200-response, 8-minute study at published corporate rates. Source: author calculation from Prolific pricing, 2026.
11Disclosure elements AAPOR requires when survey results are released publicly. Source: AAPOR Disclosure Standards.
4.2%False positive rate the authors published for the classifier behind a widely cited AI-content study. Source: Graphite, 2025.

What actually counts as original research

The short answer

Original research is any figure you produced that a reader cannot obtain from another source. It covers surveys you fielded, benchmarks you ran, and counts you made yourself from public documents. The test is not novelty of opinion. The test is whether the number still exists anywhere once your page is gone.

That definition is narrower than the one most content teams use, and I think the narrowness is the point. A roundup of other people's statistics is a convenience, not an asset. It can be replaced by any system that reads the same sources faster than your reader would.

The test is retrieval, not novelty

Ask one question about any page you are about to publish. If a language model had every source you used but not your page, could it produce the same claim? If the answer is yes, you have written a summary, and summaries are now free.

Analysis passes this test less often than writers like to believe. A sharp interpretation of public data is genuinely valuable, and it is also reproducible by anyone with the same data and a similar mind. A count, a benchmark result or a survey response is not reproducible without doing the work again.

This is the same structural argument that separates a data advantage from a workflow advantage in software, which is examined in the piece on why data moats keep collapsing into workflow moats. Publishing has the same shape. Access to the raw material is the defensible part, not the commentary layered on top.

Three formats qualify under this test: a survey you fielded, a benchmark you ran under stated conditions, and a census or audit you compiled from documents. One format does not qualify, however useful it feels. A statistics roundup assembled from other publishers is a restatement, and the analysis of how citation and rank overlap has collapsed shows how thin that ground has become.

Why summarisable content stopped working as a moat

The supply of competent, sourced, well-structured writing went from scarce to unlimited inside about two years. Graphite's classifier study found that by November 2024, more newly published web articles were primarily machine-written than human-written. The sample was roughly 43,000 URLs drawn from CommonCrawl, with publication dates between January 2020 and May 2025.

Treat that specific crossover point as directional rather than settled. The study excluded articles that were machine-drafted and then edited by a person, and its authors said they believe that category may be larger than the one they measured. The direction is not in doubt. The precise share is.

What follows is an economics problem rather than a quality problem. When the marginal cost of producing a competent summary approaches zero, the market price of a competent summary approaches zero too. The economics of the content flood works through what that does to a publishing budget.

Google's rule is about purpose, not about the tool

Publishers keep reading the scaled content abuse policy as a rule about automation. It is not. Google's spam policies define the abuse as generating many pages "for the primary purpose of manipulating search rankings and not helping users". The policy explicitly covers "large amounts of unoriginal content that provides little to no value to users, no matter how it's created".

Read the last clause carefully, because it decides the strategy. The load-bearing word is unoriginal, not automated. A pipeline that publishes 300 machine-assisted pages a month containing figures nobody else has is a different object from one publishing 300 restatements, even though both use the same tooling.

I would go further than the policy does. Original data is the only durable defence against being classified as scaled content, because it is the only property of a page that volume cannot dilute. Everything else on the checklist can be faked cheaply and at scale. The work on search visibility under content saturation covers what saturation does to the pages that lack it.

What a credible study actually costs

The most common reason small teams skip original research is a cost assumption that has not been checked since panel pricing became self-serve. The number is smaller than the assumption, and it is published.

The panel arithmetic, at published rates

Prolific lists its commercial pricing openly. The platform fee for corporate customers is 42.8% of participant rewards, charged on top of what respondents receive. Recommended participant pay is $12 per hour, with an absolute floor of $8 per hour.

From there the arithmetic is straightforward. An 8-minute survey at $12 per hour pays $1.60 per completed response. Two hundred completes cost $320 in participant pay, and the platform fee adds $136.96, for $456.96 before tax. That is the whole invoice for the fieldwork.

What a 200-response study costs to field 8-minute survey, participant pay at the recommended $12 per hour, corporate platform fee of 42.8% $320.00 Participant pay 200 x $1.60 +$136.96 Platform fee 42.8% of pay $456.96 TOTAL Invoice before VAT Excludes your time, screening rejects and design
Notice what the chart does not contain. The fieldwork is the cheap part, and every serious cost in a research programme sits outside this picture.

The costs that are not the panel

Panel fees are the least interesting line in a research budget. The expensive inputs are the ones that never appear on an invoice, and underestimating them is how research programmes die after one attempt.

Questionnaire design is the first. A badly worded question produces a clean-looking number that means nothing, and you will not discover this until someone competent reads your chart. Budget for a pilot of 20 responses before you field the real thing, and expect to rewrite at least two questions after it.

Screening is the second. If you need respondents who hold a specific job in a specific company size, your effective cost per usable response rises well above the headline figure, because you pay for screen-outs on most panels. Check the screening policy before you assume the arithmetic above holds for a narrow B2B audience.

Panel fieldwork cost at published rates, before tax and screening
Study shapeParticipant payPlatform fee at 42.8%Total
200 completes, 5 minutes$200.00$85.60$285.60
200 completes, 8 minutes$320.00$136.96$456.96
400 completes, 8 minutes$640.00$273.92$913.92
500 completes, 10 minutes$1,000.00$428.00$1,428.00

All rows are calculated from Prolific's published corporate rate of a 42.8% platform fee on participant rewards, with participant pay at the recommended $12 per hour. Figures exclude VAT, screen-outs and any incentive uplift needed to fill a narrow audience. Academic and non-profit customers pay 33.3% instead.

Sample size is a smaller obstacle than most teams assume

The second reason teams skip research is a belief that credibility starts somewhere around a thousand responses. That belief is expensive and mostly wrong, because precision improves with the square root of the sample rather than in proportion to it.

The margin of error arithmetic

For a simple random sample, the 95% margin of error on a proportion is 1.96 times the square root of p times one minus p, divided by n. At the most conservative value, where the proportion is 50%, that reduces to roughly 0.98 divided by the square root of n.

Run the numbers and the shape of the curve does the arguing. Two hundred responses gives about 6.9 points. Four hundred gives 4.9. Getting to 2.5 points needs 1,500 responses, which is more than seven times the cost of the 200-response study for less than three times the precision.

Precision improves fast, then stops improving fast 95% margin of error on a proportion at the most conservative value, simple random sample 5-point line 13.9 9.8 6.9 4.9 3.5 2.5 n=50 n=100 n=200 n=400 n=800 n=1500 The red point is the $457 study. Halving its margin of error costs four times as much fieldwork.
The curve flattens well before the sample sizes people quote as a credibility threshold. Spend the fourth dollar on better questions, not on a fifth hundred responses.

Two cautions sit alongside that arithmetic, and the second one matters more. A margin of error only describes sampling error, so it says nothing about a leading question or a panel that does not resemble your market. That formula also assumes a simple random sample, and an opt-in online panel is not one, which is exactly why the disclosure section below is not optional.

My own position is that a 200-response study with an honest limitation paragraph is more useful than a 1,000-response study presented as a fact. Readers forgive a wide interval. They do not forgive discovering that the interval was never stated.

Four research formats a small team can actually run

Not all original research is a survey, and the survey is rarely the best first project. Ranked by how defensible the output is against a determined critic, the survey comes last of the four.

The public-document audit is where I would start. Pick a population of filings, pricing pages, changelogs or job postings, define the inclusion rule in writing, and count. Cost is zero and reproducibility is total, because a sceptic can rerun your rule against the same documents and check whether they get your number.

The reproducible benchmark is second. Run a defined task under stated conditions, record the results, and publish the harness. Its weakness is version drift, since the thing you measured changes underneath the post, so a benchmark needs a dated snapshot and a rerun schedule.

Your own customer base is third. It is free, fast and genuinely proprietary, and it carries a selection problem you cannot fix. People who answer your survey are people who like you, so the results describe your customers rather than the market.

The panel survey is fourth, despite being the format everyone thinks of first. It buys you a population you do not already talk to, and that is worth real money. It also introduces panel composition questions that your customer list does not have.

Four research formats compared on cost and failure mode.
FormatCash costElapsed timeWhere it breaks
Public-document auditNone beyond your time1 to 2 weeksInconsistent disclosure across the population, which forces judgement calls you must publish.
Reproducible benchmarkCompute and API spend only1 to 3 weeksVersion drift. The result expires quietly and the post keeps ranking.
Customer-base surveyNone beyond your time2 to 3 weeksSelf-selection. Describes your customers, not the market you claim to measure.
Panel survey, 200 completesAbout $457 at published rates2 to 4 weeksPanel composition and screen-out costs on narrow B2B audiences.

The elapsed time column is editorial judgement, not measured data, and it assumes one person working part-time on the project. The cash figure in the last row is the only number in this table traceable to a published rate card. Treat the rest as a planning prior.

The disclosure block that makes a number citable

A figure without a method is a claim. A figure with a method is evidence, and the difference costs one paragraph. This is the part small teams skip, and skipping it wastes the entire budget they just spent on fieldwork.

The 11 elements, compressed into one paragraph

The American Association for Public Opinion Research publishes 11 disclosure elements for publicly released survey results. They cover the sponsor and who conducted the work, the population studied, how the sample was generated, and the modes used. The rest are the field dates, sample sizes and precision, weighting, processing rules, the instrument itself, and a statement acknowledging unmeasured error.

You do not need to print all 11 under your chart. Five of them carry almost all the credibility, and they fit in about 40 words: who paid, who was asked, how many answered, when, and what you excluded. Keep the other six on a linked methodology page for anyone who asks.

Eleven disclosure elements, two places to put them AAPOR disclosure standards for publicly released survey results PRINT THESE UNDER THE CHART 1. Who sponsored it and who conducted it 2. Population under study, defined precisely 3. Sample size and precision 4. Dates of data collection 5. Limitations and unmeasured error KEEP THESE ON A METHODOLOGY PAGE 6. Data collection strategy 7. Exact question wording and instrument 8. Sample generation, probability or not 9. Modes used to contact and collect 10. Weighting method and variables 11. Processing, validity checks, exclusions The left column is roughly 40 words. It is the whole difference between a claim and evidence.
The split is mine, not AAPOR's. AAPOR asks for all 11 on release; the argument here is only about which five earn their space directly under a chart.

Two published studies show what this looks like when it is done properly, and neither came from a research institute. The Graphite analysis named its classifier, its chunk size, its inclusion filters and its false positive rate of 4.2% against pre-ChatGPT text, then stated plainly that machine-drafted, human-edited articles were excluded.

The Content Marketing Institute did the same for its 2026 trends report, disclosing 1,015 respondents drawn from 1,229 responses, fielded from 24 June to 14 August 2025, with its three sponsors named. That report found 95% of organisations now use AI-powered applications while only 12% rate their own marketing highly effective.

Both disclosures make the studies easier to attack. That is the correct trade, and it is why they get cited instead of the dozens of unsourced statistics circulating on the same subject.

Where this argument is weakest

The case for original research gets made most often by people selling research services, usually with numbers that trace back to nothing. Here is what my own argument cannot support.

The citation evidence does not test provenance

The most quoted evidence for research content is the GEO paper from Aggarwal and colleagues. It reported visibility gains of up to 40% in generative engine responses, and it identified adding statistics as one of the effective methods. That finding is real and peer reviewed at KDD 2024.

It is also not evidence for my argument. The method inserted statistics into existing pages and measured what happened. Whether the statistics were the publisher's own was never a variable, so the paper supports statistics density and says nothing about who collected them. Anyone citing it as proof that original research wins citations, including people who agree with me, is overreading it. The review of what the citation evidence actually shows covers the wider gap between the claims and the studies.

Bad research is worse than none

A survey with a leading question or an unrepresentative panel produces a number that looks authoritative and is wrong. Publish it, get cited, and you have industrialised an error under your own name. The disclosure discipline above is what makes this recoverable, because a stated method can be corrected while a bare statistic cannot.

There is a scale limit too. One study is a post, not a moat. The advantage compounds only if the same measurement runs again on a schedule, which turns a content project into an operating commitment. Most teams underestimate that commitment, and the same measurement problem shows up in the analysis of where AI returns have actually been demonstrated.

Frequently asked questions

What is original research in content marketing?

Original research is any figure you produced that a reader cannot get from another source. It includes surveys you fielded, benchmarks you ran under stated conditions, and counts you compiled yourself from public documents. A roundup of other publishers' statistics does not qualify, because the underlying numbers exist without your page and can be restated by anything that reads the same sources.

How much does it cost to run an original research survey?

At published self-serve panel rates, about $457 buys 200 completed responses to an 8-minute survey. That figure comes from Prolific's recommended participant pay of $12 per hour plus its 42.8% corporate platform fee, and excludes VAT and screen-outs. Narrow business audiences cost more per usable response. Questionnaire design and analysis time usually exceed the fieldwork cost.

What sample size do you need for a credible industry survey?

Fewer than most people assume. At 200 responses the 95% margin of error on a proportion is roughly 6.9 points, and at 400 it is about 4.9. Precision improves with the square root of the sample, so reaching 2.5 points needs around 1,500 responses. Publishing the interval matters more than shrinking it, since an unstated margin is the actual credibility problem.

Does original research help with AI search visibility?

The honest answer is that the direct evidence is thinner than the claims. Peer-reviewed work on generative engine optimisation found adding statistics improved visibility by up to 40%, but that experiment inserted statistics into pages and never tested who collected them. Original data plausibly helps because it cannot be summarised from elsewhere, and that mechanism has not been isolated in a published study.

What should a methodology disclosure include?

AAPOR lists 11 elements for public release. Five of them belong directly under your chart: who sponsored and conducted the work, the population studied, the sample size and precision, the field dates, and a statement of limitations. The remaining six, including exact question wording, sampling method, modes, weighting and processing rules, can sit on a linked methodology page.

Is a survey of your own customers still original research?

Yes, and it carries a selection problem you should state rather than hide. Your customers are people who already chose you, so their answers describe your base rather than the market. That is still a number nobody else has. Label the population precisely as your own customers, give the response rate, and the finding remains defensible.

Where to start this month

Skip the survey. Your first project should be a public-document audit, because it costs nothing and it teaches the discipline that makes the later survey worth paying for.

Pick a population you can enumerate this week: the pricing pages of 40 competitors, every changelog entry in a category for six months, or the AI clauses in 30 published terms of service. Write the inclusion rule down before you start counting, and count the cases where the rule was ambiguous. That count of ambiguous cases is usually the most interesting finding in the whole exercise.

Then publish the five disclosure elements under the chart, in about 40 words, including the ambiguity count. If that post earns a citation, you have validated the format, and the $457 for a panel survey becomes an easy second decision rather than a speculative first one.

References

  1. Prolific, What is your pricing?, researcher help centre, read 26 August 2026. Used for the 42.8% corporate platform fee, the 33.3% academic fee, and the $12 per hour recommended participant rate. All study cost figures in this post are the author's arithmetic from those rates.
  2. American Association for Public Opinion Research, Disclosure Standards, read 26 August 2026. Used for the 11 disclosure elements and their wording.
  3. Google Search Central, Spam policies for Google web search, read 26 August 2026. Used for the scaled content abuse definition and the quoted policy text.
  4. Graphite, More articles are now created by AI than humans, 2025. Used for the sample of roughly 43,000 CommonCrawl URLs, the January 2020 to May 2025 window, the 4.2% false positive rate and the stated exclusion of human-edited AI drafts.
  5. Content Marketing Institute, MarketingProfs and Storyblok, 2026 B2B Content and Marketing Trends, October 2025. Used for the 1,015 respondents, the 24 June to 14 August 2025 field dates, the 95% AI adoption figure and the 12% highly effective figure.
  6. Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, GEO: Generative Engine Optimization, KDD 2024. Used for the up to 40% visibility figure and the list of tested methods including adding statistics.

The weakest thing about this source base is that the central claim, that first-party data earns citations that second-hand data does not, has no controlled study behind it. The GEO paper tests statistics density, not provenance. Everything here about provenance is reasoning from the mechanism, and it is labelled as such in the section on where the argument is weakest.

MJ
Aryan Vatsa
Founding Member, Zan Digital. Writes about AI product economics, B2B software markets and what the numbers behind vendor claims actually say.

Related reading