The thesis, stated
A model with world knowledge and a huge window can read my data. That is not the same as being accountable for my data. RAG exists to fix four properties that a bigger context window does not fix: membership (is this fact in the model at all), freshness (when did it last change), provenance (which document says so), and marginal cost per query. Long context moves the boundary on capacity and nothing else. Google’s own framing agrees — see the RAG Engine overview on Google Cloud — and so does the other direction of the argument, like Gemini 2.0 Flash: is this the end of RAG as we know it. The interesting part of this debate is not the slogan; it is where each side is actually correct.The limits I keep hitting
1. Private and internal data — the membership problem
Gemini does not know my company’s private database, internal HR documents, or proprietary codebases unless I feed that data to it (RAG in 5 minutes with Gemini and your data, RAG Engine overview). No amount of context window changes this: my corpus was not in the training set, and it is not in the weights. This is the limit with no workaround other than “put the data in front of the model”, which is exactly what retrieval is for.2. Freshness — the update-cost problem
If my internal database changes daily, updating a RAG vector store or using a tool like Gemini File Search is much faster than retraining or continuously uploading massive files (RAG for Gemini, The end of RAG: how Gemini File Search changes everything). Note that the second link’s title is the strongest form of the counter-argument, and it still concedes the point: File Search is retrieval, just managed by the vendor. The argument “long context ends RAG” usually turns out to be “I will let the platform run the index for me”.3. Provenance and hallucination
RAG forces the model to anchor its answers strictly to the retrieved text chunks, which drastically reduces hallucinations and provides clear source citations. Without retrieved chunks there is nothing to cite, and “trust me” is not an acceptable answer interface in a support bot or a compliance workflow. When I ship an answer with chunk IDs attached, a wrong answer becomes a debuggable event instead of a mystery.4. Cost of long context
Passing millions of tokens of raw documents into every single prompt makes API calls very slow and expensive; RAG retrieves only the exact paragraphs I need, saving time and money (Why Gemini 1.5 and other large-context models are bullish for RAG). Billing is per token per request, so a large context is not an upfront cost — it is a recurring tax on every turn of the conversation.5. Latency of long context
This is the limit people underestimate most, so I will take it apart.Can it fit is not the same as can it be fast
I ran a concrete experiment with a known small corpus: The Alchemist by Paulo Coelho, which is approximately 45,000 words long (Readingvine) — around 160 to 208 pages depending on the edition, with the standard paperback at about 208 pages, and roughly 2.5 to 4 hours of reading time for an average reader. Wikipedia’s infobox lists 163 pages for the first English edition (The Alchemist (novel)), a book first published in Brazil in 1993 in hardback, paperback and iTunes formats with an Audible audiobook, and Reading Length works out 2 hours 28 minutes at a 250-words-per-minute pace. I checked that arithmetic rather than trusting it: 45,000 words at 250 wpm is 180 minutes, which is 3 hours, not 2 hours 28 minutes. The 2h28m figure implies a shorter text than 45,000 words. Word counts for novels also vary more than people expect — Mike Shevdon notes exceptions like Stephen King’s The Gunslinger, the first Dark Tower book, at 55K words (Word-Count Worries). My takeaway: an “average document” number is a planning input with maybe 30% error on it, so I size with a range, not a point estimate. See Data size estimation for the ladder I use. Can the whole book go in one prompt? Yes, easily. The Alchemist is roughly 60,000 tokens (about 1.3 tokens per English word plus formatting — Ofox’s context window explainer), and modern models take one million to ten million tokens, so the book fits many times over. I can upload the full text to Gemini, Claude, or GPT and analyse themes, track character development, or query specific plot points all at once — Diogo Cruz documents running exactly that kind of whole-book review workflow in this narrated-review post. That is a fair, honest answer, and it is also true that at 250 wpm a human needs about three hours to do the same thing. Here is what a whole library looks like at that size:
Sources for the window sizes: LLM context window comparison 2026 (Morph), TechCompare token limits by model, Vellum’s LLM leaderboard — which lists GPT-5.6 Sol at a 1,050,000-token context — Fastio on Gemini context windows, which describes current Gemini 3 models such as Gemini 3.1 Pro and Gemini 3.8 Flash as large-input models, TechTarget’s “Gemini 1.5 Pro explained: Everything you need to know”, which compares Gemini 1.5 Pro against 1.5 Flash, ELI5: what Gemini 1.5 Pro’s 1M window means, and Gate.ai’s longest-context-LLMs survey. Context windows have expanded by roughly two orders of magnitude since the original transformer architecture, which was a few thousand tokens (Redis on LLM context windows).
Two caveats sit on top of that table. First, the window is a shared budget for input and output: if I fill it with book text, the model has no room left for a long, detailed answer (Datastudios on Gemini context strategies, Ofox). Second, fitting is not finding. Asking how effectively a model locates one specific sentence hidden inside 160 books is the “lost in the middle” question, and the answer is: worse than people assume, and worse the deeper the target sits.
The speed question, in two phases
Does it matter in terms of speed whether there is 1M or 10M tokens in context? Yes, it matters immensely. The difference can turn a near-instant answer into a wait lasting minutes (The 2M token era), and the reason is that a request has two distinct phases: reading the prompt and writing the answer (Understanding LLM response latency: input vs output processing). Phase 1 — reading, or time to first token (TTFT). Before the model writes anything it processes the whole input (Gemini support thread on high latency with large payloads, LLM latency benchmark and optimisation):- At 1M tokens the model chews through roughly 750,000 words. Prefill takes several seconds, usually 10 to 30 seconds depending on architecture and hardware (Redis, FutureAGI on Gemini 2.5 Pro).
- At 10M tokens the self-attention cost escalates dramatically — the system compares millions of tokens against each other — and prefilling a 10M-token payload easily takes 2 to 5 minutes before it begins emitting an answer (Redis, Kairi AI).
The counter-case: when RAG is genuinely not needed
I do not build a vector database reflexively. Retrieval is the wrong tool when:- The corpus fits the window and is static. One book, one contract, one 40-file module. Reading it whole beats indexing it, because chunking can split the exact paragraph where the answer lives.
- The task is global, not local. “Summarise the themes”, “compare these five candidates across twenty behavioural metrics”, “what changes between these two revisions” — these need everything in active memory at once, and top-k retrieval throws away the majority of the evidence.
- The question count is small. For a one-off analysis, the setup cost of an index never amortises. I paste the documents and pay once.
- The vendor already runs the index. Gemini File Search and managed RAG engines give me retrieval with provenance and freshness handled for me — same machinery, less of my own code.
- The latency budget is a human’s, not a product’s. A one-shot 2-to-5-minute prefill on 10M tokens is acceptable for an offline report and unacceptable for a user-facing chat turn.
The decision rule, worked on 1,000 resumes
The question I like best: I have 1,000 resumes — does it matter whether I build RAG or put everything in the context? Yes, it matters immensely. Even though a 10-million-token window can hold 1,000 resumes, RAG is vastly superior for this workload. Three pillars decide it: cost, speed, accuracy. Assume an average resume is 2 pages, roughly 1,000 words or 1,300 tokens. Then 1,000 resumes is about 1.3 million tokens of total data.
The trap is per-query billing, not the total size. Upload 1.3M tokens and ask “who has Python experience?” — I pay for 1.3M tokens. Then ask “out of those, who has worked at Google?” — I pay for another 1.3M. Cost stacks linearly with turns, which is financially unsustainable for day-to-day HR operations. The arithmetic is the same shape as the 0.7M-token cached-payload model in User usage at scale: repetition is what kills the budget, and a pipeline that re-reads the world on every turn converts a one-time cost into a per-query tax.
So my rule is short enough to put on a whiteboard:
- Corpus larger than the window, or traffic repeated across many users — retrieve. Build the index, keep it fresh, cite chunks.
- Corpus smaller than the window, question global, few repeats — load it all, cache the prefix, and skip the index.
- Both — retrieve to decide what goes in the window, then cache that window. This is the answer for most production systems I have built, and it is why
/caching-llm-chats-to-quick-answer-user-query-without-ragand this page are neighbours rather than rivals.