Ask a team about their retrieval costs and they'll usually quote the embedding price. It's a fair instinct. It's also usually the wrong line item to worry about.
This guide walks through the full economics of a retrieval-augmented generation request, from parsing a document to the last output token. You'll see which costs happen once, which repeat on every question, and where RAG cost optimization saves money without damaging relevance.

The full cost map of a RAG request

A RAG system has two cost phases. Ingestion (parsing, chunking, embedding, storing) runs once per document. Querying (retrieval, reranking, context building, generation) runs on every single question.
Ingestion costs scale with your documents. Query costs scale with your traffic.
That split explains most RAG bills. Ingestion cost grows with your corpus. Query cost grows with traffic, and in most products traffic grows a lot faster than the document collection does.
Here's what each step costs you:
StepBilled byHow often
Parsecompute time per pageonce per doc
Embedtokens embeddedonce per chunk
Storevectors and metadata heldmonthly
Retrievesearch queriesper question
Rerankcandidates scoredper question
Generateinput and output tokensper question

Why embeddings are rarely the big number

Embedding a corpus is a one-time cost per chunk. Generation input is paid on every question, so it overtakes embedding cost quickly as traffic grows, even when each question looks small.
Example: 2,000 pages at about 667 tokens each, 10,000 questions a month, 8,000 context tokens per question.
Run the numbers on a mid-sized corpus. Two thousand pages at about 667 tokens each, as Markdown, is roughly 1.33 million tokens to embed. You pay that once, plus updates.
Now serve 10,000 questions a month with 8,000 tokens of retrieved context each. That's 80 million input tokens every month. At Claude Sonnet 5.5's list price of $2 per million input tokens on the pricing page, it's about $160 a month before a single output token, and it repeats every month.
In practice, that's why the per-question context deserves your attention first. Shaving a little off every question beats a big one-time saving on ingestion.

Context size is the lever that compounds

Every token of retrieved context is billed on every question. Retrieving fewer, smaller, cleaner chunks shrinks that number directly, and the saving compounds with traffic.
Illustrative sizes. Test recall on known questions before cutting top-K.
Three settings drive it. Top-K sets how many chunks reach the model. Chunk size sets how big each one is. Cleanliness sets how much of each chunk is useful text rather than headers, footers and layout noise.
Bigger isn't automatically safer. Models handle long contexts less reliably: the 2024 paper Lost in the Middle found that information in the middle of a long input gets used less than information at the start or end. Stuffing 20 chunks in "just in case" can cost more and answer worse.

Reranking: pay a little, send a lot less

A reranker scores a wide set of retrieved candidates and keeps only the best few. It adds a cost per question, but it lets you send far fewer context tokens while improving which chunks get through.
The pattern is "retrieve wide, send narrow". Pull 20 to 50 candidates cheaply, rerank them, and pass only the top 5 to the model. In the example above, that drops context from 8,000 to 2,000 tokens per question.
There's quality evidence too. In Anthropic's 2024 Contextual Retrieval experiments, adding a reranker on top of contextual embeddings and contextual keyword search cut the top-20 retrieval failure rate by 67%, compared with 49% without it.
Source: [Anthropic, Introducing Contextual Retrieval](https://www.anthropic.com/news/contextual-retrieval), 19 September 2024.
The same article puts a price on the ingestion side. Generating a short context sentence for every chunk cost $1.02 per million document tokens, using prompt caching. For the 1.33-million-token corpus above, that's well under $2, once.

Parse quality is a cost decision

Messy parsing creates junk chunks: running headers, page numbers, split tables and broken reading order. You embed them, store them, retrieve them and pay to send them, so cleaning documents up front saves money at every later step.
A PDF converted to clean Markdown has fewer tokens per page and clearer section boundaries. That means fewer chunks to embed, and chunks that are more likely to be relevant when retrieved. Headings also make it easy to attach metadata, so you can filter by document or section before searching.
MDify converts PDFs, Office files and scans to Markdown for free. Its RAG-ready profile marks every ## heading as a chunk boundary, so your splitter can cut along sections instead of arbitrary character counts.

Where RAG cost optimization actually pays

Rank your optimizations by how often their cost repeats. Per-question costs come first, then per-document costs and storage. And before any of it, check whether you need retrieval at all.
Start at the top: those costs repeat on every question.
A practical order:
  • Set a hard context budget, such as 2,000 to 4,000 tokens, and pack the best chunks into it.
  • Retrieve wide, send narrow with a reranker.
  • Clean documents before chunking so each chunk carries more signal per token.
  • Cache the stable prompt prefix. On Claude, a cache hit costs 10% of the normal input price, per the prompt caching docs.
  • Skip RAG for small corpora. Anthropic suggests that under about 200,000 tokens, roughly 500 pages, you can often put the whole knowledge base in the prompt.
Then measure. Track RAG token usage per question, retrieval recall on a fixed test set, and answer quality. Cut context only while recall holds.

Frequently Asked Questions

What is the biggest cost in a RAG system?

For most systems with real traffic, it's the input tokens sent with each question: the retrieved context plus the prompt. Embedding is paid once per chunk, while context is paid on every question, so it grows with usage.

Are embeddings expensive?

Usually not, compared with generation. Embedding a corpus is a one-time cost per chunk, repeated only when documents change. In the example above, the corpus is 1.33 million tokens to embed once, against 80 million tokens of context sent every month.

Does reranking increase or decrease RAG cost?

Often both. It adds a per-question cost for scoring candidates, but it lets you send far fewer chunks to the model. When context tokens dominate your bill, the reranker usually pays for itself, and it can improve retrieval quality too.

How many chunks should I send to the LLM?

As few as still answer your questions reliably. Start around 5. Then test recall on a set of questions with known answers, and only add chunks when the answers need them, since more chunks cost more and can make long-context answers less accurate.

How does document cleaning reduce RAG cost?

Clean Markdown has fewer wasted tokens and clearer section boundaries. You embed fewer junk chunks, retrieve more relevant ones, and send less noise to the model on every question.

When should I not use RAG?

When the whole knowledge base is small. Anthropic suggests that below about 200,000 tokens, roughly 500 pages, you can often include everything in the prompt and use prompt caching instead.
Clean your corpus before you embed it with the free converter, and see how Markdown cuts tokens per page.