There's a habit that quietly burns most AI budgets: attach the whole file, ask one question, repeat. It works, and it's expensive every single time.
This guide lays out a practical pipeline for LLM document processing. It covers PDFs that mix text pages with scans, why Markdown is the right middle format, how to chunk with metadata, and how to assemble a context that stays inside a token budget. No framework required.
The problem with sending everything
Sending a whole document makes the model read every page for every question. You pay for all of it, and long inputs make the model worse at finding the one passage that matters.
Cost is the obvious part. A 100-page report read as text plus page images can run past 200,000 input tokens. Ask ten questions and you've paid for two million tokens of mostly irrelevant pages.
Quality is the less obvious part. The 2024 paper Lost in the Middle found that models use information at the start and end of a long context more reliably than information buried in the middle. Chroma's 2025 context rot study saw all 18 models it tested get worse as input length grew. More pages can mean worse answers, not better ones.
Step 1: extract text from every page, including scans
Extraction turns each page into text. Pages with a text layer can be read directly. Pages that are only a picture, like a scanned signature page, need text recognition first.
Real-world PDFs mix both. A contract might be generated text with two scanned annex pages. A report might include a photographed table. If your extractor treats the whole file one way, it either misses the scans or runs slow recognition on pages that already had perfect text.
The fix is per-page routing. Check each page, read text pages directly, recognize only the picture-only pages, and merge everything back in page order. That's what MDify does when you drop in a PDF: text pages are read as they are, and picture-only pages go through text recognition, up to 100 scanned pages per file.
Step 2: clean and structure it as Markdown
Cleaning removes text that repeats without adding meaning. Structuring keeps the headings, lists and tables that tell a reader, and a splitter, where each topic starts and ends.
Strip running headers, footers, page numbers and legal boilerplate. They repeat on every page, so a 40-page report can carry 40 copies of the same line into your index.
Then keep the structure as Markdown. Headings become
## lines, bullet points stay bullets, and tables become pipe tables with their header row intact. This gives the next step clean edges to cut along. It's also the format models read most cheaply, with no layout noise to wade through.MDify's RAG-ready profile goes one step further. It writes a
<!-- chunk-boundary --> marker before every ## heading, so your splitter knows exactly where sections begin.Step 3: chunk by section and attach metadata
Split the Markdown at headings so each chunk covers one topic, and give every chunk its document name, section path and page range. Metadata lets you filter before searching and cite sources after answering.
Here's a small, dependency-free Python splitter that cuts at headings, keeps the heading path, and respects MDify's chunk markers:
import re
def chunk_markdown(markdown, doc_name, max_chars=2000):
"""Split Markdown at headings and keep the heading path as metadata."""
chunks, path, lines = [], [], []
def flush():
text = "\n".join(lines).strip()
if text:
for start in range(0, len(text), max_chars):
chunks.append({
"doc": doc_name,
"section": " > ".join(path) or "(intro)",
"text": text[start:start + max_chars],
})
lines.clear()
for line in markdown.splitlines():
heading = re.match(r"^(#{1,3})\s+(.*)", line)
if heading:
flush()
level = len(heading.group(1))
path[:] = path[:level - 1] + [heading.group(2).strip()]
elif line.strip() != "<!-- chunk-boundary -->":
lines.append(line)
flush()
return chunks
On a handbook with "Refunds" and an "Exceptions" subsection, it returns chunks tagged
Handbook > Refunds and Handbook > Refunds > Exceptions. Long sections are split at max_chars so no single chunk gets out of hand.Step 4: retrieve, then assemble a context budget
At question time, search the index, drop duplicates, rerank, and pack the best chunks until you hit a fixed token budget. The budget is the step most pipelines skip, and it's the one that keeps costs flat.
Retrieval quality is worth investing in. Anthropic's 2024 Contextual Retrieval write-up reports that adding a short context sentence to each chunk, combined with keyword search, cut the top-20 retrieval failure rate by 49%. Adding a reranker took that to 67%.
Then set a hard budget, say 3,000 tokens of context. Put the strongest chunk first. If a chunk doesn't fit, leave it out rather than truncating it mid-table. A fixed budget means your cost per question stays predictable no matter how big the source documents get.
When you don't need retrieval at all
If the whole knowledge base is small, skip the index and send it all. Below a couple of hundred thousand tokens, simple can beat clever, especially with prompt caching.
Anthropic's Contextual Retrieval article makes this point directly: if your knowledge base is under 200,000 tokens, about 500 pages, you can often include all of it in the prompt. With prompt caching, a cache hit on that repeated prefix costs 10% of the normal input price.
Even then, convert to Markdown first. A 500-page corpus as clean Markdown is far smaller than the same corpus as PDF pages with images, so it fits more easily and caches more cheaply.
| Corpus size | A sensible approach |
|---|---|
| A few pages | Paste the relevant section as Markdown |
| Under ~500 pages | Whole corpus as Markdown, cached |
| Larger, or growing | Chunk, index, retrieve within a budget |
Frequently Asked Questions
What is a document to LLM pipeline?
It's the set of steps that turns raw files into input a model can use efficiently. You extract text, clean it, structure it as Markdown, split it into chunks with metadata, index it, and retrieve only the relevant chunks per question. The heavy work happens once per document, not once per question.
Why not just send the whole PDF to the AI?
Because you pay for every page on every question, and long inputs can reduce answer quality. Studies like Lost in the Middle (2024) and Chroma's context rot research (2025) show models handle long contexts less reliably. Sending the right few sections is cheaper and often more accurate.
Why convert documents to Markdown before chunking?
Markdown keeps headings, lists and tables as plain text, so a splitter can cut at section boundaries and each chunk stays on one topic. It also drops the page images and layout noise that make PDFs expensive to read.
How big should chunks be?
There's no single right size. Start with one section per chunk, a few hundred tokens each, and split long sections. Then test with questions you already know the answers to, because the best size depends on your documents and your search method.
Can this pipeline handle scanned PDFs?
Yes, if the extraction step routes pages. Read pages with a text layer directly and run text recognition only on picture-only pages. MDify does this automatically for up to 100 scanned pages per PDF, then merges the pages back in order.
Do I need a vector database for this?
Not always. For a small corpus, sending everything as cached Markdown can be simpler. Once your documents grow past a few hundred pages, or change often, an index with retrieval keeps each question's cost flat.
Start with step one: convert your documents with the free converter, then see how Markdown saves tokens in practice.