Skip to main content
This is the complete pipeline, built twice: once as a version you can run in the next five minutes with nothing installed but Python, and once as the version I would actually put in front of users. The reason I insist on the dependency-free first pass is that almost every RAG explanation I have read starts with a vector database and an API key, which means the reader cannot see the retrieval working, cannot tell which stage failed, and cannot debug the one that did. I ran the local version in this environment and pasted its real output below, including the parts where it does badly. Environment: Python 3.14.6, standard library only — numpy, sentence_transformers, qdrant_client, openai, and fastapi were all absent, which is exactly the constraint the fallback path is designed for. Anything that needed a package or a key is labelled and given as an exact command. The companion video for this topic is in the references at the bottom, and the architecture around it is in RAGs.

What a RAG pipeline is, in stages

1

Chunk

Split documents into retrieval units of roughly 250 to 400 tokens. The unit is what I retrieve, so its boundaries decide what the model can ever see.
2

Embed or index

Turn each chunk into something searchable — a vector for semantic search, an inverted index for lexical search. Both are retrieval; only one needs a model.
3

Retrieve

Score the query against the index, take top-k, and keep the provenance (document id, chunk id) attached.
4

Prompt

Assemble a context block from retrieved chunks with citation anchors, then the instruction and the question.
5

Generate

Ask the model to answer only from that context and to cite the anchors it used.
6

Evaluate

Score retrieval (recall, precision, MRR) and generation (groundedness, answer correctness) separately, because they fail differently and cost differently to fix.

The runnable version, end to end

Lexical BM25 stands in for the vector index, and a deterministic extractive selector stands in for the LLM. Everything else is a real RAG pipeline. Save as minirag.py and run python3 minirag.py.

Verified output

This is what python3 minirag.py printed in my environment, not a hand-written approximation:
Two things in that output are worth pausing on. The corpus is 8 documents and 13 chunks with an average chunk length of 16.08 terms, and the top hit is a 7.76-score gap over the second — that is BM25 doing exactly what it should on a discriminative term. And answer-match=0.40 against answer-in-context=1.00, which I will explain in a moment because it is the most instructive number on the page.

The retrieval score, worked by hand

For the term bellamar in chunk d1#0, measured from the same run:
The structure is idf × tf-saturation. df=3 out of N=13 chunks gives idf = ln(1 + (13 − 3 + 0.5)/(3 + 0.5)) = 1.3863. The saturation term with k1=1.2, b=0.75, tf=1, chunk length 21 and avgdl=16.08 is (1 × 2.2) / (1 + 1.2 × (1 − 0.75 + 0.75 × 21/16.08)) = 0.8887. Their product, 1.2320, is that term’s contribution to the chunk’s score. Two properties I want you to see: a term appearing more times saturates toward k1 + 1 rather than growing linearly, and a chunk longer than average is penalised by b. Those are the two knobs that explain most “why is this result ranked there” conversations.

Reading the evaluation honestly

Metric arithmetic for this run, so you can check my numbers: ranks were 1, 1, 2, 1, 1, so MRR = (1 + 1 + 0.5 + 1 + 1) / 5 = 0.900 and recall@3 = 5/5 = 1.00. Per-query doc-level precision was 1/3, 1/3, 1/3, 2/3, 1/3 — the fourth query got two chunks from the same gold document inside top-3 — so precision@3 = 2.0/5 = 0.40. Now the interesting part. Retrieval was near-perfect and generation still failed on 3 of 5 questions. I deliberately printed answer-in-context alongside answer-match to make that separation visible: the correct answer sentence was present in the retrieved context for all 5 questions, and my generator only got 2 of them. The clearest case is how much does the airport transfer cost. Traced from the same run:
The right document was retrieved, and both its chunks were in the top-3, which is why recall is 1.00. But the chunk carrying the price ranked below the chunk carrying the booking rule, because transfer does not match transfers, and cost has no lexical bridge to fee or euro — df(cost) = 0, so it contributes nothing at all. Three fixable failure modes in one query: no stemming or lemmatisation, no synonym handling, and no semantic signal. That is exactly the query-document mismatch and chunking-loses-context pair from the FDE interview syllabus, and it is why the production path below is not optional in a real system.
Do not ship an aggregate score. A pipeline with recall@3 = 1.00 looks flawless and still answers 60% of questions wrong. Split retrieval metrics from generation metrics, and print both every run, or you will debug the wrong stage for a week.

The production path, per stage

Each stage above has a real replacement. I did not run these in this environment because the packages and keys are absent, so treat the commands as the exact ones to execute.
Then evaluate with a framework rather than my hand-rolled loop, which is what the stack in Sample Resume uses: pip install ragas, score faithfulness (is the claim supported by context), answer relevancy, context precision, and context recall — the last two being exactly the split my table above makes with precision@3 and answer-in-context.
Order of investment in a real system, based on what actually moved metrics for me: chunking quality first (a badly bounded chunk cannot be retrieved well by any index), then hybrid retrieval with fusion, then the reranker, then the prompt, then the model swap. People buy the last one first and wonder why the answers are still wrong.
The size of your context block is a budget decision, not a habit. Five chunks of 300 tokens plus the question is roughly 1,550 input tokens per query — see Data size estimation for how that number becomes a monthly invoice, and Caching LLM Chats to Quickly Answer User Queries Without RAG for when to skip retrieval and serve from a cached prefix instead. For the index those chunks live in, Scaling a Vector Database and this case study.

What I would change next, in order of payoff

Having the pipeline runnable means each of these is an experiment I can score rather than a guess. Chunk size. My max_words=25 chunks are deliberately small, and small chunks retrieve precisely but answer poorly, because the sentence that answers “what time is check-in” loses the sentence two lines later that says it applies only to loyalty members. Large chunks do the opposite: fewer vectors, cheaper index, more context per hit, and worse ranking because a chunk about five things matches a query about one of them. I sweep 128, 256, and 512 tokens and read recall@k and answer-match together — the pair always disagrees, and the disagreement is the finding. Overlap. Zero overlap is a coin flip on boundary-spanning answers. Adding a sentence of overlap costs vectors — the 1.2x multiplier in Data size estimation — and buys back the boundary cases. It is the cheapest quality improvement available and the first thing I measure. Top-k. With k=3 I already have recall@3 = 1.00 on this corpus, so a larger k would only add tokens and cost. On a real corpus the opposite holds: raise k for recall, then let the reranker cut it back down before the prompt. The reranker is how I get both a wide candidate pool and a small context budget. Stemming, then semantics. The transfer versus transfers miss is fixed by stemming, which is roughly twenty lines and stdlib-free in this design. The cost versus fee miss cannot be fixed lexically at all — df(cost) = 0 means the term contributes nothing — and that is precisely the gap dense embeddings exist to close. Once a real embedding model is in place I keep BM25 alongside it and fuse with RRF, because the lexical channel is what pins exact identifiers, prices, and error codes. Prompt and abstention. My instruction already says “if the context does not answer the question, say so”, and the offline generator implements abstention when the score is zero. That behaviour is not a nicety: in production the most expensive failure is a confident answer over an empty or wrong context, and I want an explicit, measurable abstain rate rather than a hallucination rate discovered by a customer.

Cost and latency of what I just built

The toy corpus hides the economics, so I size them. This pipeline retrieves 5 chunks of about 300 tokens, roughly 1,550 input tokens per query before instructions. Against the Why Need RAG if Gemini Can Answer alternative — putting the entire corpus in context on every turn — retrieval is not a quality decision first, it is a per-query billing decision, and the difference grows with every follow-up question in the same session. So the pipeline has a fixed cost I amortise and a variable cost I have to budget per turn, and the tuning knobs above move the variable line directly.
Three failure modes this run demonstrates on purpose, all of them on the standard interview syllabus in What to Expect in FDE GenAI Roles: chunking losing context (boundary effects), query-document mismatch (transfer versus transfers, cost versus fee), and stale indexes (my corpus is a dict literal; a real one must be re-embedded when the document or the model version changes). Recall versus precision is the trade I made visible by printing both.

References