What is RAG (Retrieval Augmented Generation)
The Problem RAG Solves
Large Language Models are trained on vast amounts of text up to a specific cutoff date. This creates two critical limitations for business applications:
Knowledge cutoff: The model does not know about events, documents, or data created after its training cutoff. If your business's products, policies, or knowledge base changed after the model was trained, the model will give outdated or incorrect answers.
Hallucination: When asked about specific facts they were not trained on -- like your company's internal pricing, a specific contract clause, or a customer's account history -- LLMs often generate plausible-sounding but incorrect information.
Retrieval Augmented Generation (RAG) solves both problems by giving the model access to relevant, up-to-date information at the time of each query.
What is RAG?
RAG is an architecture where the AI system retrieves relevant documents from a knowledge base before generating a response. Instead of relying solely on its training data, the model uses both its knowledge and the retrieved documents to produce accurate, grounded answers.
The term was introduced in a 2020 paper by Facebook AI Research, and has since become one of the most widely deployed AI architectures in production.
How RAG Works: Step by Step
User question: "What is our refund policy for digital products?"
Step 1: ENCODE the question
Convert the question into a vector embedding (a list of numbers representing its meaning)
Step 2: RETRIEVE relevant documents
Search the vector database for documents semantically similar to the question
Result: ["returns-policy.pdf page 3", "customer-faq.md section 5", "terms-and-conditions.txt"]
Step 3: AUGMENT the prompt
Include the retrieved documents in the prompt to the LLM:
"Using only the following documents, answer the question: [retrieved documents] Question: [user question]"
Step 4: GENERATE the response
The LLM generates an answer grounded in the retrieved documents, not speculation
The retrieval step is what makes RAG different from a simple LLM call. The model is not guessing -- it is reading relevant documentation and answering based on what it finds.
Why RAG is Better than Fine-Tuning for Most Use Cases
When developers first encounter the knowledge cutoff problem, they often consider fine-tuning the model on their data. RAG is usually the better choice:
| Factor | RAG | Fine-Tuning |
|---|---|---|
| Cost | Low (embedding + inference) | Very high (GPU training costs) |
| Data updates | Real-time -- add new docs instantly | Requires retraining |
| Transparency | Sources are visible and citable | Model just "knows" things |
| Control | Easy to add/remove documents | Difficult to remove learned knowledge |
| Hallucination risk | Low -- grounded in retrieved docs | Higher -- model may confuse training data |
The only time fine-tuning is clearly preferable is when you want to change the model's style, tone, or format capabilities -- not when you want to add factual knowledge.
RAG in Practice: Real-World Applications
Customer Support Knowledge Base
A company has 500 support articles, 10,000 product FAQs, and 200 policy documents. Instead of hiring 50 more support agents, they build a RAG system:
- Embed all documents once
- When a support query arrives, retrieve the top 3-5 relevant documents
- Ask the LLM to answer based only on those documents
- The system cites the specific article it used
Result: 70% of tier-1 support queries answered automatically, with source citations for every answer.
Legal Document Q&A
A law firm uploads hundreds of case files and legal documents. Lawyers ask questions like "What are the termination conditions in the Apex Corp contract?" The RAG system retrieves the specific contract section and answers accurately -- without the lawyer having to manually search 100+ pages.
Internal Knowledge Base
A company's internal policies, procedures, and technical documentation are embedded. New employees ask questions like "What is the expense approval process?" and get accurate answers from the latest policy documents.
Personalised Financial Advice
A bank uses RAG to give customers personalised answers: the customer's account data, transaction history, and relevant banking regulations are retrieved and used to ground the response.
The Components of a RAG System
- Document store: Where original documents are stored (files, databases, wikis)
- Embedding model: Converts text into vector representations
- Vector database: Stores and indexes embeddings for fast similarity search
- Retrieval engine: Given a query, finds the most semantically similar documents
- LLM: Generates the final answer based on the retrieved context
- Response formatter: Extracts the answer and optionally cites sources
Naive RAG vs. Advanced RAG
The simple version described above is called Naive RAG. Production systems often use advanced techniques:
Chunking strategy: How you split documents into pieces affects retrieval quality. A 500-word chunk is often more useful than splitting by sentence or by entire document.
Hybrid search: Combining semantic search (vector similarity) with keyword search (BM25) often outperforms either alone.
Re-ranking: After retrieving 20 candidate chunks, use a re-ranking model to select the 5 most relevant.
Query transformation: Expand or rephrase the user's query to improve retrieval. "Tell me about returns" becomes ["return policy", "refund process", "product returns FAQ"].
Multi-hop RAG: For complex questions, retrieve once, identify what additional information is needed, then retrieve again.
Key Takeaways
- RAG solves the LLM knowledge cutoff and hallucination problems by retrieving relevant documents before generating a response.
- The four RAG steps are: Encode the query, Retrieve relevant documents, Augment the prompt with those documents, Generate the answer.
- RAG is almost always preferable to fine-tuning for knowledge-based tasks -- it is cheaper, updatable in real time, and produces transparent, citable answers.
- Common applications include customer support Q&A, legal document analysis, internal knowledge bases, and personalised advisory systems.
- Production RAG systems improve basic retrieval with techniques like hybrid search, re-ranking, and query transformation.
Try it yourself
Key Takeaways
- RAG (Retrieval Augmented Generation) solves LLM knowledge cutoff and hallucination by retrieving relevant documents before generating answers.
- The four RAG steps are: encode the query, retrieve similar documents, augment the prompt with those documents, generate a grounded answer.
- RAG is almost always preferable to fine-tuning for domain knowledge -- it is cheaper, real-time updatable, and produces transparent, citable answers.
- Common RAG applications include customer support Q&A, legal document analysis, internal knowledge bases, and personalised advisory systems.
- Production RAG systems improve retrieval with hybrid search, re-ranking, query transformation, and optimal chunk sizing.
Quick Quiz
1.What does RAG stand for and what problem does it solve?
2.Why is RAG usually preferable to fine-tuning for adding domain knowledge to an LLM?
3.What are the four steps in the RAG pipeline?
4.What is 'hybrid search' in the context of advanced RAG systems?
Ready to go further?
CareerEx gives you structured 12-week training, live classes every Saturday and Sunday, real tutor feedback, and a certificate. Join the next cohort.
Join CareerEx