From Shubhi K | Product & Market Analysis
RAG vs Long Context: We Priced a 10,000-Document Knowledge Base Both Ways
On this page
A 10,000-document knowledge base does not fit in a 1M-token context window. At an assumed 2,000 tokens a document it is 20 million tokens, so reading all of it per question costs $40.13 on Claude Sonnet 5.5 at list price. Retrieval augmented generation answers the same question for under 3 cents. On cost, RAG vs long context is not a close contest.
Key takeaways
- A 1M-token window holds about 5% of a 10,000-document base. At 2,000 tokens a document the base is 20 million tokens, so "just put it all in context" means 20 full-window calls per question.
- Uncached, the full-base approach costs about 1,460 times more per query than RAG. On Claude Sonnet 5.5 list rates, that is $40.13 against $0.0275, before any vector database fee.
- Caching narrows the gap but does not close it. With cache reads at $0.10 per million tokens, reading the whole base still costs $2.13 a query, and Gemini charges $4.50 per million tokens per hour just to keep a cache alive.
- The real break-even is corpus size, not query volume. On Sonnet 5.5 a cached corpus costs the same per query as a 10,000-token RAG prompt at 200,000 tokens, roughly 100 documents. Above that, retrieval wins.
The short answer
Long context is cheaper than RAG only when your whole corpus is small. On Claude Sonnet 5.5, a cached corpus under about 200,000 tokens costs the same or less per query than a 10,000-token retrieval prompt. A 10,000-document knowledge base is 100 times bigger than that, so retrieval wins at every query volume.
The 10,000-document knowledge base we priced
Most RAG vs long context comparisons price a single long prompt against a single short one. That framing quietly assumes your knowledge base fits in the window. For a real company knowledge base, it usually does not.
So this model starts from the corpus, not the prompt. Picture a support and policy library of 10,000 documents: help articles, contracts, runbooks, product specs. That is a modest base for a mid-sized software company.
The assumptions, stated
Every figure below rests on five assumptions. They are ours, not a vendor's, and you should replace them with your own before trusting any total.
| Input | Value used | Why this value |
|---|---|---|
| Average document length | 2,000 tokens | Roughly a four-page help article. Your mix may run longer. |
| Corpus size | 20 million tokens | 10,000 documents multiplied by 2,000 tokens. |
| RAG prompt per question | 10,000 tokens | Instructions, the question, and 10 retrieved chunks of about 900 tokens. |
| Answer length | 500 output tokens | A full paragraph answer with citations. |
| Reference model | Claude Sonnet 5.5 | $2 input, $10 output, $0.10 cache read, $4 one-hour cache write, per million tokens. |
Rates are list prices from Anthropic's pricing page as read on 9 October 2026. Month length is 30 days. No volume discounts, no Batch API, no data residency uplift.
Why a 1M-token window is not enough
Claude Sonnet 5.5, Opus 5.5 and Haiku 5.5 all carry a 1M-token context window, per Anthropic's context window documentation. Gemini 3.1 Pro and OpenAI's GPT-6.1 Sol sit in the same range.
A 1M window therefore holds about 500 of your 10,000 documents. To answer from the whole base without retrieval, you must split it into 20 shards, send the question to each, and merge the answers. That is the long-context architecture this post prices, because it is the only one that genuinely skips retrieval.
The alternative people actually build is one window holding a hand-picked 5% of the base. That is cheaper, and it is also a retrieval system. Someone, or something, chose which 500 documents went in.
Long context pricing on three official rate cards
The vendors price long prompts very differently. Anthropic dropped its long-context premium in March 2026. OpenAI and Google still charge more once a prompt crosses a threshold.
| Model | Surcharge starts at | 1M tokens, uncached | 1M tokens, cached read | Same model, 10K-token RAG prompt |
|---|---|---|---|---|
| Claude Sonnet 5.5 | No surcharge. | $2.00 | $0.10 | $0.020 |
| Claude Haiku 5.5 | Over 100,000 tokens. | $0.50 | $0.05 | $0.001 |
| OpenAI GPT-6.1 Sol | Over 272,000 tokens. | $4.00 | $0.20 | $0.020 |
| Gemini 3.1 Pro Preview | Over 200,000 tokens. | $4.00 | $0.40, plus $4.50 per hour of storage | $0.020 |
Sources: Anthropic, OpenAI and Google API pricing pages, read 9 October 2026. Above each threshold the whole request is billed at the higher rate. The final column uses each model's below-threshold rate, which is why it does not scale with the others.
Where the surcharges start
The thresholds matter more than the headline rates. On OpenAI's pricing page, GPT-6.1 Sol doubles its input rate from $2 to $4 above 272,000 tokens, and output rises from $10 to $15. Gemini 3.1 Pro doubles from $2 to $4 above 200,000 tokens.
Claude Haiku 5.5 is the odd case. It is 5 times dearer above 100,000 tokens, yet still the cheapest 1M prompt on the table at $0.50. Anthropic states that a 900k-token request is billed at the same per-token rate as a 9k-token request on its other current models.
Notice the final column. Below the threshold, all three flagship models charge the same $0.02 for a 10,000-token RAG prompt. Retrieval keeps you in the cheap tier on every vendor, which is a pricing advantage nobody advertises.
The RAG cost model at three query volumes
Here is the same knowledge base priced four ways at 1,000, 10,000 and 100,000 questions a month. Every rate traces to a vendor pricing page. Every multiplication is shown in the assumptions above.
| Architecture | Per query | 1,000 queries | 10,000 queries | 100,000 queries | Coverage |
|---|---|---|---|---|---|
| Full base, 20 windows, uncached | $40.13 | $40,125 | $401,250 | $4,012,500 | All 10,000 documents. |
| Full base, 20 windows, cached | $2.13 plus writes | $4,525 | $23,650 | $214,900 | All 10,000 documents. |
| One 1M window, cached | $0.105 plus writes | $225 | $1,170 | $10,620 | About 500 documents. |
| RAG with reranker and vector store | $0.0275 plus $50 | $77.50 | $325 | $2,800 | All 10,000, searched. |
Cached rows assume one full cache write per day at the one-hour write rate: $80 a day for the full base, $4 a day for one window. Low traffic with gaps over an hour will force more writes. RAG includes a Voyage rerank-3 call at its published estimate of $0.0025 a request and Pinecone's $50 Standard minimum.
What RAG actually costs to run
The retrieval side is cheaper than most teams assume. Embedding the entire 20M-token base with voyage-4-lite or OpenAI's text-embedding-3-small costs $0.40 at $0.02 per million tokens. Voyage's current models include 200 million free tokens, so the first full embed can cost nothing.
Storage is similarly small. At about 900 tokens a chunk, the base becomes roughly 22,000 vectors, a fraction of a gigabyte. Pinecone's Standard plan has a $50 monthly minimum and a $20 flat Builder tier, and pgvector on a Postgres you already run adds close to nothing. The full comparison is in the Pinecone, Weaviate and pgvector cost breakdown.
The generation call dominates. On Sonnet 5.5, 10,000 input tokens cost $0.02 and 500 output tokens cost $0.005. At 100,000 questions a month, model tokens are $2,500 of the $2,800 total. Our position is simple: spend your optimisation effort on the prompt, not the vector store.
Prompt caching changes the long context answer, but only for small corpora
Prompt caching is the strongest argument for long context. You pay once to write a large prefix, then pay a fraction to reread it. On Claude Sonnet 5.5, a cache read costs 0.05 times the base input price, so 1M cached tokens cost $0.10 instead of $2.
That 20-fold discount is real, and it is why the cached full-base row falls from $401,250 to $23,650 at 10,000 queries. It is still 73 times the RAG total, because reading 20 million cached tokens per question is still reading 20 million tokens.
The break-even rule for RAG vs long context
The useful question is not "which is cheaper". It is "at what corpus size does caching beat retrieval". The arithmetic is short. A cached corpus costs its size times the cache rate. A RAG prompt costs its size times the base rate.
They match when the corpus equals the RAG prompt divided by the cache multiplier. On Sonnet 5.5 that is 10,000 tokens divided by 0.05, or 200,000 tokens. On models with the standard 0.1 multiplier, such as Haiku 4.5 or Sonnet 4.6, it is 100,000 tokens.
Below 200,000 tokens, we would skip the retrieval stack entirely. You lose a moving part, a failure mode and an index to keep fresh. Above that line, every document you add raises the per-question cost of long context and leaves RAG untouched.
Two vendors make caching less generous than it looks. Gemini 3.1 Pro bills storage at $4.50 per million tokens per hour, so caching the full 20M-token base costs $64,800 a month in storage alone. OpenAI's cached reads on GPT-6.1 Sol double to $0.20 per million above 272,000 tokens. The wider pattern, where a falling unit price still produces a rising bill, is covered in the analysis of why cheap tokens keep raising AI bills.
The accuracy bill arrives after the invoice
Cost is only half the trade. The other half is whether the model actually uses what you paid to put in front of it. On that question, the published evidence is unkind to very long prompts.
The NoLiMa benchmark, published at ICML 2025, tested 13 models that claim at least 128K tokens of context. It hid facts whose wording barely overlapped the question, so models had to infer the link rather than match keywords. At 32K tokens, 11 of the 13 fell below half their own short-context score. GPT-4o dropped from 99.3% to 69.7%.
Newer models do better, and the vendors say so. Anthropic reported in March 2026 that Opus 4.6 scored 78.3% on MRCR v2 at 1M tokens, the highest among frontier models at that length. Read the other side of that number: the best reported model still missed more than a fifth of retrievals at full window.
That is the hidden cost of stuffing. You pay $40 for a question, and you buy a model attending to 20 million tokens with known blind spots. We covered the diagnosis side in the guide to fixing bad RAG before abandoning it, and the newer discipline of choosing what goes in the window in the piece on context engineering.
Where the case for RAG is weakest
This post argues for retrieval on cost. Here is where that argument loses, stated as plainly as the rest.
Long context wins on quality when you can afford it
The best controlled comparison we found favours long context on accuracy. Li and colleagues at Google DeepMind and the University of Michigan tested three models and concluded that, when resourced sufficiently, long context consistently outperforms RAG on average. The study is from 2024 and used models two generations old, which limits how far it transfers.
Their Table 1 shows the size of the gap. Gemini-1.5-Pro scored 49.70 with long context and 37.33 with RAG; GPT-4o scored 48.67 against 32.60. The same paper found RAG and long context gave identical answers on 63% of queries, which is the opening for a hybrid.
Be careful with these figures. The paper's own text and tables disagree in places, for example quoting a 65% cost cut for Gemini where Table 1 implies about 62%. We use the table values throughout.
The engineering cost this model leaves out
Our RAG column prices tokens and storage. It does not price the engineer who builds chunking, tunes retrieval, maintains an evaluation set and debugs a missed document at 2am. For a small team, that salary can dwarf a few thousand dollars of tokens.
Two smaller caveats cut the same way. Anthropic notes its newer tokenizer produces about 30% more tokens for the same text, so your base may be 26 million tokens on Claude, not 20. And our 2,000-token document is an assumption; a base of short FAQ entries could fit in one window.
A RAG architecture for 2026 that uses both
The framing of RAG or long context is the wrong question. The cost model says retrieval should decide what enters the window, and the large window should give retrieval room to be generous.
In practice, that means a three-step design. Retrieve 30 to 50 candidate chunks, rerank them, and send the top 20 or so into a prompt of 20,000 to 40,000 tokens. That prompt stays below every vendor's surcharge threshold and costs $0.04 to $0.08 in input on Sonnet 5.5.
Then add a router for the hard cases. When the model reports low confidence, escalate that one question to a cached long-context call over the relevant document set. Most questions take the cheap path, and the expensive path is reserved for the minority that need it, which is the pattern model routing applies to model choice.
Finally, measure it. A retrieval system you cannot score is one you cannot tune, and building an evaluation suite is the cheapest insurance in this whole design. Without it, you will not know whether a missed answer came from retrieval or from the model.
Frequently asked questions
Is RAG cheaper than long context?
For any corpus larger than a few hundred thousand tokens, yes. On Claude Sonnet 5.5 list rates, reading a 20-million-token knowledge base through 1M windows costs about $40 per question uncached and about $2 cached. Retrieving 10,000 tokens of relevant chunks costs about 3 cents including a reranker. Long context only wins on cost when the whole corpus is small enough to cache cheaply.
Does a 1 million token context window replace RAG?
Not for most company knowledge bases. A 1M window holds about 500 documents of 2,000 tokens, so a 10,000-document base needs 20 windows per question. Benchmarks also show accuracy falls as context grows, and the best vendor-reported model still missed over a fifth of retrievals at full length. A large window makes RAG more forgiving, it does not remove the need for it.
How much does it cost to run RAG on 10,000 documents?
Using Claude Sonnet 5.5 at list price, about $325 a month at 10,000 questions and $2,800 at 100,000. That assumes 10,000 input tokens and 500 output tokens per answer, a reranker call, and a $50 vector database minimum. Embedding the full 20-million-token base once costs about 40 cents at $0.02 per million tokens.
When should I use long context instead of RAG?
Use long context when your full corpus is under about 200,000 tokens on Claude Sonnet 5.5, or 100,000 tokens on models with a 0.1 cache multiplier. At that size, a cached prompt costs the same as a retrieval prompt and removes an index to maintain. Also use it for tasks that need reasoning across a whole document, such as reviewing one long contract.
Does prompt caching make long context cheaper than RAG?
Caching cuts the cost of rereading a prompt by 90% or more, depending on the model. It does not change how many tokens you read per question. Caching a 20-million-token base on Sonnet 5.5 still costs about $2 per question, roughly 77 times a RAG call. Gemini also charges hourly storage, which adds $64,800 a month to keep that base cached.
Which LLMs charge more for long context prompts?
As of October 2026, OpenAI's GPT-6.1 Sol doubles its input rate above 272,000 tokens and Gemini 3.1 Pro doubles above 200,000 tokens. Claude Haiku 5.5 charges 5 times more above 100,000 tokens. Anthropic's other current Claude models bill the full 1M window at standard rates, with no long-context surcharge.
Where to start this week
Measure your corpus before you choose an architecture. Run your actual document set through your target model's token counter and write the total down. If it is under 200,000 tokens, try a cached prompt first and keep the vector database off the roadmap.
If it is larger, take your current monthly question volume and multiply it by the per-query figures in the table above. Then compare that total with one engineer-month. Whichever number is bigger tells you where your next optimisation hour belongs.
References
- Anthropic, Claude API pricing, read 9 October 2026. Used for all Claude token, cache, long-context and tokenizer figures.
- Anthropic, Context windows documentation, read 9 October 2026. Used for 1M window availability by model.
- OpenAI, API pricing, read 9 October 2026. Used for GPT-6.1 Sol rates, the 272K threshold and embedding prices.
- Google, Gemini Developer API pricing, last updated 7 October 2026. Used for Gemini 3.1 Pro rates and cache storage.
- Voyage AI, Pricing, updated 29 September 2026; and Pinecone, Pricing, read 9 October 2026. Used for embedding, reranking and vector storage costs.
- Modarressi et al., NoLiMa: Long-Context Evaluation Beyond Literal Matching, ICML 2025. Used for the 32K accuracy findings.
- Li et al., Retrieval Augmented Generation or Long-Context LLMs? A Comprehensive Study and Hybrid Approach, July 2024. Used for LC vs RAG scores and Self-Route token ratios.
- Anthropic, 1M context is now generally available for Opus 4.6 and Sonnet 4.6, 13 March 2026. Used for the MRCR v2 score.
The weakest part of this source base is the accuracy evidence. The strongest controlled LC vs RAG study predates every model priced here, and the most recent long-context score is vendor-reported. Every cost figure is a model built on our stated assumptions, not a measured invoice.
Related reading