· 8 min
Markdown vs Raw PDFs: Which Uses Fewer LLM Tokens?
A PDF page can cost an AI app three times the tokens of the same page as Markdown. Here's where that gap comes from, and when sending the PDF is still the right call.
Guides to clean Markdown, document conversion and AI token costs.
· 8 min
A PDF page can cost an AI app three times the tokens of the same page as Markdown. Here's where that gap comes from, and when sending the PDF is still the right call.
· 7 min
Switching to a cheaper model is the least interesting way to save money on LLM APIs. Here are seven levers that cut waste instead, with the math behind each one.
· 7 min
The PDF file isn't the problem. What happens when an AI reads it is. Here's where the extra tokens come from, page by page, and how to stop paying for them.
· 7 min
Sending a 100-page PDF to answer one question pays for 99 pages you didn't need. Here's the pipeline that sends only the part that answers.
· 6 min
Most RAG pipelines send five times the context the model needs, and nobody notices because the answers still look fine. Here's where that token tax comes from.
· 6 min
Teams worry about embedding prices while the real RAG bill grows at question time. Here's the full cost map, with the numbers that show where to optimize.
· 7 min
Your questions are usually the smallest part of what you send an AI. Here are the six places tokens quietly pile up, and a 30-minute audit to find yours.
· 6 min
Most LLM bills aren't high because AI is expensive. They're high because each request carries tokens nobody needed. Here's how one question grows to 31,970 tokens.