What Is RAG? Retrieval-Augmented Generation Explained (with Python Code)
Ask a chatbot about your company's refund policy, your team's internal wiki or last week's meeting notes, and it has no idea. It was never trained on your documents, and even if it were, it might confidently invent an answer. Retrieval-augmented generation, or RAG, fixes both problems, and it's the most common way businesses put large language models to work on their own data.
This post explains how RAG works step by step, with diagrams, then builds a tiny working version in Python. It builds on how large language models work, especially the ideas of the context window and hallucination.
The problem RAG solves
A large language model knows only what was in its training data. That leaves three big gaps:
- Private knowledge: it has never seen your documents, emails, tickets or databases.
- Fresh knowledge: anything after its training cutoff is missing.
- Trust: when it doesn't know, it may produce a plausible but wrong answer, with no source to check.
RAG's idea is refreshingly simple: before the model answers, look up the most relevant passages in your documents and hand them to the model along with the question. It's the difference between a closed-book exam and an open-book exam.
RAG hands the model the right pages from your documents before it answers.
How RAG works: two phases
A RAG system has two parts. Indexing prepares your documents once, in advance. Answering runs every time someone asks a question.
The whole RAG system in two phases. Indexing happens once; the answering loop runs for every question.
Step 1: Split documents into chunks
You can't paste a 40-page manual into every prompt: it's slow, expensive and may not fit in the context window. Instead, documents are split into chunks, typically a few hundred words each. Chunks usually overlap slightly so an idea that spans a boundary isn't lost.
Documents are split into chunks small enough to search precisely and fit into the prompt.
Chunk size matters. Too small, and a chunk lacks the context to be useful. Too large, and search becomes imprecise and prompts fill up with irrelevant text. Splitting along natural boundaries, such as headings and paragraphs, usually works better than cutting at a fixed character count.
Step 2: Turn chunks into embeddings
Each chunk is converted into an embedding: a list of numbers that captures its meaning, produced by an embedding model. Texts with similar meanings get similar numbers, even when they use different words.
Embeddings place text by meaning. A question about getting “money back” lands next to chunks about refunds, even without shared words.
This is what makes RAG smarter than a keyword search. "How do I get my money back?" and "Refund policy" share no words, but their embeddings sit close together. The embeddings are stored in a vector database, such as FAISS, Chroma, Pinecone or PostgreSQL with the pgvector extension, which can find the nearest neighbors among millions of vectors in milliseconds.
Step 3: Retrieve the most relevant chunks
When a question arrives, it's embedded with the same model, and the vector database returns the chunks whose embeddings are closest, typically the top 3 to 10. Closeness is usually measured with cosine similarity, a score where 1 means pointing in exactly the same direction.
Step 4: Build the prompt and generate
The retrieved chunks are pasted into a prompt, together with instructions and the user's question:
The final prompt: clear instructions, the retrieved chunks as context, and the user's question.
The instruction to answer only from the context, and to admit when the answer isn't there, is what keeps the model grounded. Many systems also ask the model to cite which chunk each fact came from, so users can check the source.
Build a mini RAG in Python
The program below implements indexing, retrieval and prompt building for a small set of customer-support documents. To keep it runnable anywhere with no downloads or API keys, it uses TF-IDF, a classic keyword-based technique from scikit-learn, in place of a neural embedding model. The final step, sending the prompt to an LLM, is shown as commented-out code, because it needs an API key.
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.metrics.pairwise import cosine_similarity
# 1. Your documents (in a real app: PDFs, wiki pages, tickets, manuals...)
documents = [
"Refund policy: customers can request a full refund within 30 days of purchase. "
"After 30 days, refunds are issued as store credit.",
"Shipping: standard orders ship within 2 business days. "
"Express shipping delivers in 1 to 2 days for an extra fee.",
"Warranty: all devices include a 1-year limited warranty covering manufacturing defects. "
"Accidental damage is not covered.",
"Support hours: the help desk is open Monday to Friday, 9am to 6pm Eastern time.",
"Account security: enable two-factor authentication in Settings under Security.",
]
# 2. Index: turn every document into a vector (here TF-IDF; real apps use embeddings)
vectorizer = TfidfVectorizer(stop_words="english")
doc_vectors = vectorizer.fit_transform(documents)
# 3. Retrieve: find the documents most similar to the question
def retrieve(question, k=2):
q_vector = vectorizer.transform([question])
scores = cosine_similarity(q_vector, doc_vectors)[0]
best = scores.argsort()[::-1][:k]
return [(round(float(scores[i]), 2), documents[i]) for i in best]
# 4. Augment: put the retrieved text into the prompt
def build_prompt(question):
hits = retrieve(question)
context = "\n".join(f"- {text}" for _, text in hits)
return (
"Answer the question using only the context below. "
"If the answer is not in the context, say you don't know.\n\n"
f"Context:\n{context}\n\nQuestion: {question}"
), hits
for question in ["Is accidental damage covered by the warranty?",
"Can I get my money back after 45 days?"]:
print("Q:", question)
for score, text in retrieve(question):
print(f" score={score} {text[:55]}...")
prompt, _ = build_prompt("Can I get my money back after 45 days?")
print("\n" + prompt)
# 5. Generate: send the prompt to any LLM. For example, with the Anthropic SDK
# (pip install anthropic, and set the ANTHROPIC_API_KEY environment variable):
#
# import anthropic
# client = anthropic.Anthropic()
# reply = client.messages.create(
# model="claude-sonnet-5", max_tokens=300,
# messages=[{"role": "user", "content": prompt}],
# )
# print(reply.content[0].text)
Running it prints:
Q: Is accidental damage covered by the warranty?
score=0.67 Warranty: all devices include a 1-year limited warranty...
score=0.0 Account security: enable two-factor authentication in S...
Q: Can I get my money back after 45 days?
score=0.42 Shipping: standard orders ship within 2 business days. ...
score=0.37 Refund policy: customers can request a full refund with...
Answer the question using only the context below. If the answer is not in the context, say you don't know.
Context:
- Shipping: standard orders ship within 2 business days. Express shipping delivers in 1 to 2 days for an extra fee.
- Refund policy: customers can request a full refund within 30 days of purchase. After 30 days, refunds are issued as store credit.
Question: Can I get my money back after 45 days?
The first question works perfectly: the warranty document scores highest by a wide margin. The second result is more instructive. Keyword matching ranked the shipping document first, simply because both it and the question contain the word "days", while "money back" and "refund" share no words at all. The refund policy still made it into the prompt only because we retrieved two chunks.
This is exactly why production RAG systems use neural embeddings instead of keyword matching: an embedding model understands that "money back" means "refund". Given this context, a good LLM can still answer correctly: after 30 days, refunds are issued as store credit. To upgrade the example, replace the TF-IDF lines with an embedding model, such as one from the sentence-transformers library, and keep everything else the same.
Making RAG work well in practice
- Clean your documents. Remove boilerplate, fix broken text from PDFs, and keep titles and headings with each chunk. Bad input is the most common cause of bad answers.
- Use hybrid search. Combine embeddings with keyword search. Embeddings capture meaning; keywords are better for exact terms like product codes and names.
- Re-rank the results. Retrieve 20 candidates quickly, then use a more careful re-ranking model to pick the best 5.
- Add metadata filters. Restrict search by date, department or document type, and respect user permissions so people only retrieve what they're allowed to see.
- Evaluate systematically. Build a set of test questions with known answers, and measure both retrieval (was the right chunk found?) and the final answer (was it correct and grounded?).
RAG versus fine-tuning
| RAG | Fine-tuning | |
|---|---|---|
| Best for | Answering from specific, changing knowledge | Teaching a style, format or specialized skill |
| Updating knowledge | Add or edit documents, instantly | Retrain the model |
| Shows sources | Yes, it can cite retrieved chunks | No |
| Cost to start | Low | Higher: needs training data and compute |
For most "chat with our documents" projects, RAG is the right first step. The two can also be combined.
Key terms, decoded
- RAG (retrieval-augmented generation)
- Retrieving relevant text and adding it to the prompt before an LLM generates its answer.
- Chunk
- A small piece of a document, sized for search and for fitting into a prompt.
- Embedding
- A list of numbers representing the meaning of a piece of text.
- Vector database
- A database built to find the stored embeddings most similar to a query embedding.
- Cosine similarity
- A score measuring how closely two embeddings point in the same direction.
- Grounding
- Making the model base its answer on supplied sources rather than its memory.
- Hybrid search
- Combining meaning-based and keyword-based search.
- Re-ranking
- A second, more careful pass that reorders retrieved results by relevance.
RAG turns a general-purpose LLM into an assistant that knows your material: split, embed, retrieve, then generate. It's simple enough to prototype in an afternoon, and it's the foundation of most real-world AI assistants. A follow-up post will build a complete RAG app with real embeddings and a vector database, end to end.
What documents would you most like to chat with? Tell us in the comments.

Comments
Post a Comment