[ AI Academy ]
Retrieval-Augmented Generation (RAG)
Master Retrieval-Augmented Generation (RAG) with comprehensive tutorials, Python code examples, and interactive exercises
← Back to Learning HubRetrieval-Augmented Generation (RAG)
Retrieval-Augmented Generation (RAG) is a technique that enhances Large Language Models (LLMs) by combining retrieval and generation. Instead of relying solely on a model's training data, RAG systems retrieve relevant documents or knowledge from external sources and use them to augment the generation process, producing more accurate, grounded, and up-to-date responses.
The RAG approach solves critical limitations of standalone LLMs: hallucinations (generating false information), lack of access to real-time or domain-specific knowledge, and inability to trace where information comes from. By integrating a retrieval component, RAG systems can provide citations, handle private documents, and adapt to knowledge updates without retraining.
RAG has become essential in production systems for question-answering, customer support, research assistance, and knowledge work. Companies like OpenAI, Anthropic, and every major AI company have adopted RAG-style approaches in their products. Understanding RAG is critical for anyone building practical AI applications.
What You'll Learn
RAG Architecture
Understand the complete RAG pipeline: document ingestion, chunking, embedding, storage, retrieval, and generation.
Retrieval Systems
Learn vector databases (FAISS, Chroma, Pinecone), embedding models, and retrieval metrics that power modern RAG.
Advanced Techniques
Explore re-ranking, query expansion, multi-hop reasoning, and hybrid search to build production-grade RAG systems.
Prerequisites
Familiarity with Python, LLMs (GPT, BERT embeddings), basic machine learning concepts, and ideally some experience with vector databases or information retrieval.
Why RAG Matters
LLMs are powerful but flawed. They hallucinate, have fixed knowledge cutoffs, can't access proprietary documents, and can't update their knowledge without retraining. RAG solves these problems elegantly by separating knowledge retrieval from knowledge generation.
The Problem with LLMs Alone
Hallucinations
LLMs generate plausible-sounding but false information. A 7B parameter model hallucinates ~28% of the time on factual questions.
Stale Knowledge
Models trained in 2023 don't know about 2024 events. Retraining is expensive and impractical for frequent updates.
No Private Data
LLMs can't access your company's documents, databases, or internal knowledge bases without fine-tuning.
No Attribution
Users can't verify where information came from. In regulated industries, this is unacceptable.
How RAG Solves It
RAG elegantly addresses all four problems:
RAG significantly improves accuracy, attribution, and knowledge currency
RAG in Production
Why companies use RAG: OpenAI uses RAG in ChatGPT's browsing mode. Anthropic uses it for retrieval. Google uses it in Search Generative Experience. Every enterprise AI system uses RAG to ground models in proprietary data and reduce hallucinations.
Historical Evolution of RAG
RAG didn't emerge fully formed in 2020. It represents the convergence of information retrieval research (decades old) with modern deep learning and LLMs.
The Evolution Timeline
BM25 and TF-IDF: Traditional keyword matching dominated. These algorithms ranked documents by term frequency and inverse document frequency. Fast, interpretable, but limited to exact keyword matches.
Vector Space Models: Documents represented as vectors; similarity computed via cosine distance. This foundational idea still underlies modern RAG.
Word2Vec (2013): Dense word embeddings showed that semantic relationships could be captured in vector space. This was revolutionary.
Neural Ranking: Researchers began using neural networks to learn ranking functions instead of hand-crafted BM25 weights.
BERT (2018): Contextual embeddings that understand meaning in context, not just individual words.
BiDAF and Reading Comprehension: Models that could read documents and answer questions about them. This was the precursor to RAG.
Dense Passage Retrieval (DPR, 2020): Facebook AI Research showed that learning dense embeddings specifically for retrieval was more effective than BM25. This was crucial โ it showed retrieval itself could be learned end-to-end.
RAG Paper (Lewis et al., Facebook, 2020): "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" formally defined RAG. The key insight: condition generation on retrieved documents.
Vector Databases Become Practical: Pinecone (2021), Weaviate, and others made vector databases accessible. No longer required specialized expertise.
LLMs Arrive: GPT-3 (2020) showed that large models could do amazing things. But they still hallucinated.
LangChain (2022): Made RAG accessible to everyone with simple APIs for retrieval, generation, and tool use.
Production RAG at Scale: Companies deployed RAG for customer support, internal search, research, and knowledge work.
Advanced RAG: Re-ranking, query expansion, hypothetical document embeddings, multi-hop reasoning, and more.
RAG Evaluation (RAGAS, 2023): Frameworks for measuring RAG quality. No longer just eyeballing results.
Agentic RAG (2024): Systems that decide when to retrieve, what to retrieve, and iterate on results. RAG becomes a tool in a broader reasoning system.
Core Concepts
To understand RAG, you need to understand a few foundational concepts that come from information retrieval and machine learning.
Embeddings and Semantic Search
An embedding is a numerical representation of text in a high-dimensional space (typically 384-1536 dimensions). Semantically similar texts have embeddings close in vector space.
Measures angle between vectors. Range: -1 to 1. Higher is more similar.
Embedding Models convert text to vectors:
- BERT / RoBERTa: Bi-directional context. Slow but accurate. Use for important tasks.
- Sentence-Transformers: Optimized for semantic search. Fast and effective. Industry standard.
- OpenAI text-embedding-3-small/large: Proprietary, powerful, but costs add up at scale.
- Specialized Models: BGE, E5, ColBERT for domain-specific or multi-lingual retrieval.
Vector Databases
Storing and searching embeddings efficiently requires specialized databases optimized for similarity search:
- FAISS (Facebook): In-memory, blazing fast, used when you fit data in RAM.
- Chroma: Simple, developer-friendly, good for prototypes and small-to-medium applications.
- Pinecone: Fully managed, scalable, built for production with metadata filtering.
- Milvus / Weaviate: Open-source, scalable, support complex queries.
- Elasticsearch with vector capabilities: For teams that already use Elasticsearch.
The RAG Pipeline: Core Steps
RAG systems follow a standard flow:
RAG Pipeline Flow
Key Terms
- Chunking: Breaking documents into overlapping pieces (typically 256-1024 tokens) to manage context length and improve retrieval granularity.
- Relevance: Whether retrieved documents actually answer the query. Measured by similarity score and human judgment.
- Ranking: Ordering retrieved results by relevance. Naive RAG ranks by embedding similarity; advanced RAG uses cross-encoders to rerank.
- Context Window: The maximum tokens the LLM can see. Modern LLMs have 4K-200K tokens; RAG must fit retrieved docs + query + instructions in this budget.
- Hallucination: LLM generating false information not grounded in retrieved documents. RAG reduces but doesn't eliminate hallucination.
RAG Architecture & Pipeline
A complete RAG system consists of multiple stages. Understanding each stage and how they interact is critical to building effective systems.
The Full RAG Architecture
RAG System Architecture
Stage 1: Document Ingestion
Raw documents come from many sources: PDFs, websites, databases, APIs, internal wikis, research papers, product docs. The ingestion system must handle various formats and scale to process millions of documents.
Document Sources
Common sources include customer support docs, research papers, internal knowledge bases, product documentation, website content, and real-time data feeds.
Stage 2: Preprocessing & Chunking
Documents must be broken into manageable pieces (chunks). Too small chunks lose context; too large chunks exceed context windows. Typical chunk size: 256-1024 tokens with 20-50% overlap.
chunk_size = 512
overlap_size = chunk_size ร overlap_pct
= 154 tokens overlap between chunks
Stage 3: Embedding
Convert text chunks to dense vectors using embedding models. This happens offline once; at query time, queries are embedded in the same space.
Stage 4: Storage & Indexing
Vectors are stored in vector databases with metadata (source document, chunk ID, date, etc.). These databases use specialized indexes (HNSW, IVF, PQ) for fast similarity search.
Stage 5: Query Processing
User query is embedded using the same embedding model. This is critical โ using different embedding models for documents and queries breaks the system.
Stage 6: Retrieval
Find top-K similar documents (typically K=3-10) using approximate nearest neighbor search. Trade-off between speed and accuracy.
Stage 7 (Optional): Re-ranking
Use a more powerful model (cross-encoder) to re-rank retrieved results. Slower but more accurate. Optional but recommended for high-quality systems.
Stage 8: Generation
Combine query + retrieved context + system prompt, and feed to LLM. LLM generates answer grounded in retrieved documents.
Key Components of RAG Systems
RAG systems have several critical components. Understanding and optimizing each is key to building production systems.
1. Embedding Models
Sentence-Transformers
Fast, open-source models like all-MiniLM-L6-v2 (384 dims). Great for general-purpose semantic search.
OpenAI Embeddings
Proprietary text-embedding-3 (1536 dims). Powerful but costs increase with scale.
Domain-Specific
BGE for multilingual, ColBERT for specialized domains. Trade speed for accuracy.
Sparse vs Dense
Dense embeddings capture semantics; sparse (BM25) capture exact keywords. Hybrid often works best.
2. Chunking Strategies
Fixed-Size Chunking
Split documents into fixed-size pieces (e.g., 512 tokens). Simple but loses structure. Risk of splitting sentences mid-way.
Semantic Chunking
Use sentence or paragraph boundaries. Preserves meaning but variable size can complicate batch processing.
Recursive Chunking
Split by hierarchy: sentences โ paragraphs โ sections. Preserves semantic structure. Recommended approach.
3. Vector Databases
Choose based on your scale and infrastructure:
4. Retrieval Methods
- Semantic Search: Query and documents in same embedding space. Foundation of RAG.
- Hybrid Search: Combine dense vectors + sparse BM25. Excellent results, more computation.
- Metadata Filtering: Retrieve by category, date, or other metadata before similarity search. Reduces hallucination.
- Multi-hop Retrieval: Iterative retrieval: retrieve โ answer intermediate question โ retrieve again. For complex queries.
- Query Expansion: Generate multiple versions of query (synonyms, rephrasings) and retrieve for all. Increases recall.
5. Re-ranking (Critical for Quality)
Embedding-based retrieval is fast but can miss nuance. Cross-encoder re-rankers improve relevance:
scored = cross_encoder.predict(
[(query, doc) for doc in retrieved_docs]
)
reranked = sorted_by_score(scored)[:10]
Popular Re-rankers
ms-marco-MiniLM for speed, ms-marco-TinyBERT for extreme speed, mmarco-mMiniLMv2 for multilingual. Each trades accuracy vs latency.
6. Context Management
LLMs have context limits. You must fit: [system prompt] + [retrieved docs] + [query] + [generation space] within limit.
- Context Pruning: Summarize retrieved docs or keep only most relevant parts.
- Sliding Window: If too much context, keep top-K most relevant chunks and summarize the rest.
- Long-Context Models: GPT-4 Turbo (128K), Claude 3 (200K). Higher costs but more flexibility.
Implementation Guide: Building a Basic RAG System
Let's build a working RAG system from scratch. We'll use popular libraries: LangChain for orchestration, Sentence-Transformers for embeddings, and Chroma for vector storage.
Step 1: Install Dependencies
Step 2: Load Documents
Step 3: Create Embeddings and Store in Vector DB
Step 4: Create a Retriever
Step 5: Set Up LLM and Build RAG Chain
Step 6: Add Source Attribution
Advanced RAG Techniques
Basic RAG is powerful, but production systems need advanced techniques to achieve high quality. Here are the most important optimizations.
1. Re-ranking with Cross-Encoders
Initial retrieval via embeddings is fast but imprecise. Cross-encoders compute interaction scores between query and document, providing more accurate relevance scores.
2. Query Expansion
Generate multiple versions of the query to improve recall. Especially useful for complex or ambiguous questions.
3. Hypothetical Document Embeddings (HyDE)
Instead of embedding the query directly, generate a hypothetical answer, embed that, then use it to retrieve. This works surprisingly well!
4. Multi-hop Retrieval (Iterative RAG)
For complex questions requiring multiple steps, retrieve โ answer intermediate question โ retrieve again.
5. Hybrid Search (Dense + Sparse)
Combine semantic search (dense vectors) with keyword search (BM25). Hybrid often outperforms either alone.
6. Metadata Filtering
Filter retrieval by metadata (source, date, category) before similarity search. Reduces hallucinations.
Comparing RAG Approaches
There are multiple ways to implement RAG. Each has tradeoffs in complexity, quality, latency, and cost.
Naive RAG vs Advanced RAG vs Agentic RAG
When to Use Each
Naive RAG
Use for MVP/prototypes, real-time applications where latency is critical, or simple fact retrieval. Good for: chatbots answering FAQs, simple document search.
Advanced RAG
Use for production systems where quality matters, enterprise knowledge bases, research assistants. Good for: customer support, internal search, documentation QA.
Agentic RAG
Use for complex reasoning, multi-step problem solving, when decisions about what/when to retrieve are important. Good for: research assistance, complex analysis, reasoning-heavy tasks.
Real-World RAG Use Cases
RAG is transforming how organizations interact with information. Here are the most impactful applications.
Customer Support & Q&A
Support Ticket Automation
RAG searches support docs, FAQs, and past tickets to suggest or auto-generate responses. Reduces response time by 50-80%, reduces support costs significantly.
Interactive FAQ
Users ask questions in natural language; RAG retrieves relevant FAQ entries and generates comprehensive answers with citations.
Chatbot Personalization
RAG retrieves customer history, account details, and relevant docs to provide personalized support. Higher satisfaction scores.
Enterprise Knowledge Management
Internal Search
Employees search across documentation, wikis, and policies. RAG understands intent and returns relevant information with context.
Research Assistant
Scientists and analysts use RAG to search papers, data, and findings. Helps synthesize information across thousands of documents.
Policy Compliance
RAG helps identify relevant policies and regulations for queries. Important for legal, finance, and regulated industries.
Content & Creative
Content Generation
RAG retrieves source material (research, interviews, data) to ground content generation. Reduces fabrication, improves accuracy.
Code Search
Developers search codebases using natural language. RAG finds relevant code snippets and documentation. GitHub Copilot uses this.
Product Recommendations
RAG retrieves product info, reviews, and specs. Generates personalized recommendations with explanations.
Healthcare & Regulatory
Clinical Decision Support
RAG retrieves medical literature, treatment guidelines, and patient history. Assists doctors in decision making. Requires high accuracy.
Drug Discovery
Researchers use RAG to search papers on compounds, mechanisms, and trials. Accelerates identification of research directions.
Regulatory Compliance
Retrieve relevant regulations, requirements, and precedents. Critical for healthcare, finance, and pharma.
E-commerce & Marketplace
Smart Product Search
RAG understands shopper intent and retrieves relevant products from detailed catalogs. Improves conversion by 15-30%.
Seller Support
Marketplace sellers use RAG to search policies, shipping rules, and best practices. Reduces policy violations.
Enterprise RAG Applications
Enterprise RAG systems must handle scale, security, compliance, and integration with existing systems. This section covers enterprise-grade patterns.
Scalability Patterns
Multi-Collection Architecture
Store different data sources in separate vector collections. Allows different embedding models, refresh rates, and metadata schemas. Query relevant collections based on intent.
Sharding
For massive document collections (100M+), shard across multiple vector databases. Route queries to relevant shards. Requires careful metadata strategy.
Hybrid Deployment
Mix cloud-based vector DBs (Pinecone) with on-premises or private deployments for sensitive data. Use federated search.
Security & Privacy
- Data Isolation: Separate collections per tenant/organization. Ensure queries only retrieve authorized documents.
- Encryption: Encrypt documents at rest and in transit. Vector DBs support encrypted storage.
- Access Control: Implement row-level security. Tag documents with access levels and filter during retrieval.
- Audit Logging: Log all queries and retrievals for compliance and debugging.
- PII Handling: Anonymize or mask sensitive data in retrieved documents before sending to LLM.
Quality Assurance
Evaluation Metrics
Track retrieval accuracy (recall, precision), generation quality (factuality, relevance), and end-to-end user satisfaction.
Feedback Loops
Collect user feedback on answer quality. Use feedback to identify failing queries and improve chunking/retrieval strategies.
A/B Testing
Test different embeddings, chunk sizes, retrieval strategies, and LLMs. Measure impact on quality and latency.
Integration with Enterprise Systems
Data Connectors
Integrate with Salesforce, Jira, Confluence, SharePoint, databases. Keep vector store in sync with source data.
API Gateway
Expose RAG system via API. Implement rate limiting, authentication, and request logging.
Workflow Integration
Embed RAG into existing workflows. Trigger retrieval from document management systems, CRM, or chat platforms.
Cost Management
- Cache Retrieved Documents: Store frequently retrieved docs in cache. Reduces redundant embedding computations.
- Batch Embedding: Embed documents in batches during off-peak hours, not in real-time.
- Model Selection: Use smaller, faster embedding models (384-dim) instead of large ones when possible. Fine-tune on domain data for better quality without scaling to larger models.
- Selective Re-ranking: Only re-rank top-20 results, not all results.
- Query Optimization: Cache embeddings of common queries. Deduplicate similar queries before retrieving.
Common RAG Mistakes (And How to Avoid Them)
Building RAG systems is deceptively simple but deploying high-quality production systems is hard. Here are the most common pitfalls.
1. Using Different Embedding Models for Indexing vs Querying
Mistake: Index with OpenAI embeddings but query with Sentence-Transformers. Why Bad: Vectors are in different spaces. Similarity scores are meaningless. Fix: Use the same model for indexing and querying, always.
2. Ignoring Chunk Overlap and Context
Mistake: Split documents into 256-token chunks with 0% overlap. Why Bad: Important concepts get split across chunks, retrieved documents lack context. Fix: Use 20-50% overlap. Test different chunk sizes.
3. Not Evaluating Retrieval Quality
Mistake: Only evaluate end-to-end answer quality. Why Bad: Can't diagnose if problems are in retrieval or generation. Fix: Evaluate retrieval metrics (recall, precision, NDCG) separately from generation quality (factuality, relevance).
4. Over-Optimizing for Latency
Mistake: Use small embeddings, no re-ranking, minimal retrieval to reduce latency. Why Bad: Answer quality suffers. Fix: Measure quality-latency tradeoff. Users prefer accurate over fast.
5. Forgetting About LLM Hallucinations Even With RAG
Mistake: Assume retrieved documents prevent hallucinations. Why Bad: LLMs still hallucinate even when given correct context. Happen ~10-20% of the time even with good retrieval. Fix: Implement fact-checking, confidence scoring, and require citations.
6. Not Handling Updates to Documents
Mistake: Embed documents once and never update. Why Bad: System becomes stale. Critical for real-time or frequently updated data. Fix: Implement incremental updates. For some data, re-embed daily or hourly.
7. Using Raw Documents Without Metadata
Mistake: Store chunks with no metadata (source, date, category, access level). Why Bad: Can't filter, can't trace results, violates compliance. Fix: Rich metadata on every chunk. Use for filtering and attribution.
8. Ignoring the Cost of Embeddings at Scale
Mistake: Use expensive embedding APIs without considering cost. Why Bad: At 1M documents, cost becomes prohibitive. Fix: Use open-source models (Sentence-Transformers). Only use proprietary APIs when necessary.
9. Not Testing on Your Actual Data
Mistake: Test RAG on toy datasets, deploy on real data. Why Bad: Real data has different characteristics. Quality may be much lower. Fix: Evaluate on representative samples of your actual data early and often.
10. Underestimating the Importance of Query Understanding
Mistake: Embed query as-is without preprocessing. Why Bad: Typos, abbreviations, and poor phrasing hurt retrieval. Fix: Implement query expansion, spelling correction, and intent classification before retrieval.
RAG Best Practices
Based on what works in production, here are proven best practices for building high-quality RAG systems.
Document Preparation
- Clean Data: Remove headers, footers, boilerplate. Preserve semantic structure. Quality in โ quality out.
- Rich Metadata: Attach source, date, category, confidence, access level, version. Use for filtering and attribution.
- Optimal Chunking: 256-1024 tokens is typical. Test different sizes. Recursive chunking > fixed-size chunking.
- Overlap: 20-50% overlap between chunks. Helps maintain context at chunk boundaries.
Embedding Selection
- Start with Sentence-Transformers: all-MiniLM-L6-v2 (384-dim) is a strong baseline. Free, fast, effective.
- Fine-tune if needed: If domain-specific vocabulary is important, fine-tune an embedding model on your data.
- Multilingual: If serving multiple languages, use multilingual models (multilingual-e5-small).
- Consistency: Same model for indexing and querying. Document which model was used.
Retrieval Strategy
- Start Simple: Semantic search (dense retrieval) works well. Add complexity only if needed.
- Hybrid Search for Recall: Combine dense + sparse (BM25). Higher recall with minimal latency increase.
- Always Re-rank: If quality is critical, re-rank top-20 with cross-encoder. ~500ms extra latency, 10-20% quality improvement.
- Metadata Filters: Filter by date, source, category before similarity search. Reduces noise and hallucination.
Generation
- Context Window Management: Fit [prompt + retrieved docs + generation] in context. Prioritize retrieved docs, summarize if needed.
- System Prompt: Clear instructions to use retrieved docs, cite sources, say "I don't know" if not in docs.
- Temperature: Low temperature (0.2-0.3) for factual tasks. Higher (0.7+) for creative tasks.
- Stopping Tokens: Use stop tokens if generating citations in specific format.
Evaluation & Monitoring
- Retrieval Metrics: Recall@K, Precision@K, NDCG. Essential for diagnosing issues.
- Generation Metrics: Use RAGAS (Retrieval-Augmented Generation Assessment) framework. Measures answer relevance, groundedness, and citation accuracy.
- User Feedback: Thumbs up/down on answers. Use to identify and fix failing queries.
- Continuous Testing: Build evaluation pipeline. Test weekly on new data.
Deployment
- Start Small: Begin with small dataset, expand after validating quality.
- Version Control: Version embeddings and chunking strategy. A/B test new versions.
- Monitoring: Track latency, error rates, cost. Alert on quality degradation.
- User Experience: Show confidence scores, source documents, alternative answers.
Advanced Insights & Future Directions
RAG is rapidly evolving. Here are emerging techniques and future directions shaping the field.
Sparse-Dense Hybrids
Recent research shows that hybrid retrieval (combining sparse BM25 with dense vectors) often beats either alone. More nuanced: different retrieval methods excel at different aspects (density captures semantics, sparsity captures keywords). Future systems will intelligently choose methods based on query type.
Self-Improving RAG
Systems that learn from retrieval failures. When answer is wrong, identify whether fault was in retrieval or generation. Adapt: better query expansion for retrieval failures, better prompts for generation failures.
Multimodal RAG
Extending RAG beyond text: retrieve images, tables, code snippets, videos. Require multimodal embeddings (CLIP, LLaVA) and multimodal generation models. Frontier is still being mapped.
Graph-based RAG
Instead of flat document retrieval, build knowledge graphs from documents. Retrieve subgraphs relevant to query. Allows reasoning across multiple documents and relationships. More complex but higher quality for knowledge-intensive tasks.
Long-Context Models Changing the Game
GPT-4 Turbo and Claude 3 have 100K-200K context windows. Some researchers are experimenting with putting entire corpora in context (in-context RAG). Extreme approach: index everything in the prompt. Reduces retrieval quality concerns but increases cost and latency.
Agentic RAG
RAG as a tool in reasoning systems. LLM decides what to retrieve, when, and how many times. Learns to reason about information gaps. Early results show promise but current LLMs struggle with this level of metacognition.
Fine-Tuning on Retrieval Tasks
Instruction-tuned models specifically for retrieval-augmented generation tasks. E.g., training models to generate better queries, to recognize when they need to retrieve, or to synthesize information from multiple docs.
Latency Optimization
Current bottleneck: embedding the query and retrieving from vector DB (100-200ms). Future: cached embeddings, streaming retrieval, early stopping. Goal: sub-50ms retrieval for real-time applications.
Complete Code Examples
Working, production-ready code for building RAG systems. Copy-paste and run.
Example 1: Basic RAG with LangChain
Example 2: RAG with Re-ranking
Example 3: Advanced RAG with Query Expansion and Filtering
Example 4: RAG Evaluation with RAGAS
Hands-On Exercises
Learning RAG is best done by building. Here are practical exercises to deepen your understanding.
Exercise 1: Build Your First RAG System
Create a Basic RAG Pipeline
Load a PDF document (download any research paper from arXiv), chunk it, embed it, and build a RAG system that answers questions about it. Steps: 1. Find a research paper PDF (aim for 5-15 pages) 2. Use PyPDFLoader to load it 3. Chunk with RecursiveCharacterTextSplitter (512 tokens, 20% overlap) 4. Create embeddings with Sentence-Transformers 5. Store in Chroma 6. Build QA chain with LangChain 7. Test with 5-10 questions about the paper Success Criteria: System answers at least 80% of questions correctly with proper citations.
Exercise 2: Compare Embedding Models
Evaluate Different Embeddings
Take the same documents and create vector stores with 3 different embedding models. Compare retrieval quality on the same test queries. Models to compare: - all-MiniLM-L6-v2 (fast, small) - all-mpnet-base-v2 (balanced) - all-distilroberta-v1 (small, decent) Metrics: - Retrieval latency - Relevance (manual scoring 1-5) - Memory usage What you'll learn: How to choose embeddings for different requirements.
Exercise 3: Advanced Retrieval with Re-ranking
Implement Cross-Encoder Re-ranking
Add a re-ranking layer to your RAG system. Compare quality before/after re-ranking. Setup: 1. Use your basic RAG from Exercise 1 2. Retrieve top-20 documents (instead of top-5) 3. Re-rank with ms-marco-MiniLM cross-encoder 4. Keep top-5 re-ranked results Evaluation: - Does re-ranking improve answer quality? - How much latency cost? - Is it worth it for your use case?
Exercise 4: Implement Query Expansion
Build Query Expansion
Implement query expansion: given a query, generate 3 alternative phrasings and retrieve from all. Compare to single-query retrieval. Process: 1. Use an LLM to expand queries 2. Retrieve for original + expanded 3. Deduplicate results Measure: - Recall improvement - Latency increase
Exercise 5: Evaluate with RAGAS
Setup RAG Evaluation
Install RAGAS and evaluate your RAG system on 20 test questions. Metrics to track: - Faithfulness (does answer follow context?) - Answer relevancy (does it answer the question?) - Context precision (are retrieved docs relevant?) Goal: Achieve >0.85 on all metrics.
RAG Interview Questions
Preparing for interviews? Here are the questions you're likely to face on RAG systems.
Foundational Questions
Architecture & Design Questions
Technical Implementation
Quality & Advanced Topics
Frequently Asked Questions
Summary
RAG is the practical solution for building AI systems that are accurate, grounded, and up-to-date. Here's what you've learned:
Key Takeaways
RAG Fundamentals
RAG combines retrieval (finding relevant docs) and generation (LLM creating answers). Solves hallucination, stale knowledge, and attribution problems.
Complete Pipeline
Document ingestion โ chunking โ embedding โ storage โ query processing โ retrieval โ generation. Each stage matters.
Production Patterns
Naive RAG is simple but limited. Advanced RAG (re-ranking, expansion) gives 15-20% quality boost. Choose based on requirements.
Evaluation
Measure retrieval quality (recall, precision) separately from generation (faithfulness, relevancy). Use RAGAS framework.
Next Steps
- Build a basic RAG system using the code examples in this guide.
- Evaluate on your data using RAGAS or similar frameworks.
- Optimize retrieval with re-ranking and query expansion.
- Deploy to production with monitoring and user feedback loops.
- Stay updated โ RAG is evolving rapidly with new techniques emerging monthly.
Common Pitfalls to Avoid
- Different embedding models for indexing vs querying
- Not evaluating retrieval quality separately
- Expecting RAG to completely eliminate hallucinations
- Ignoring updates to source documents
- Not tracking costs at scale
Resources for Continued Learning
- Papers: "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al.) โ the seminal RAG paper
- Libraries: LangChain, LlamaIndex for RAG orchestration
- Evaluation: RAGAS framework for systematic evaluation
- Communities: LangChain Discord, RAG research papers on arXiv
Resources & Further Learning
Essential Papers
- "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al., 2020) โ The original RAG paper. Must read.
- "Dense Passage Retrieval for Open-Domain Question Answering" (Karpukhin et al., 2020) โ Foundational work on learned dense retrieval.
- "Realm: Retrieval-Augmented Language Model Pre-Training" (Guu et al., 2020) โ Pre-training with retrieval.
- "In-Context Retrieval-Augmented Language Models" (Ram et al., 2023) โ Modern RAG patterns with in-context learning.
- "RAGAS: A Unified Metric for Evaluating Retrieval-Augmented Generation" (Es et al., 2023) โ How to properly evaluate RAG.
Tools & Libraries
- LangChain (langchain.com) โ Python library for RAG orchestration. Most popular.
- LlamaIndex (llamaindex.ai) โ Alternative RAG framework, strong for document indexing.
- Sentence-Transformers (sbert.net) โ Open-source embedding models.
- Chroma (trychroma.com) โ Vector database for development.
- Pinecone (pinecone.io) โ Managed vector database for production.
- RAGAS (github.com/explodinggradients/ragas) โ Evaluation framework for RAG systems.
- DSPy โ Formal framework for composing RAG programs.
Learning Resources
- LangChain Documentation โ Best practices and tutorials
- Hugging Face Hub โ Thousands of pre-trained embedding models
- arXiv.org โ Latest research papers on retrieval and generation
- GitHub โ See real implementations in production projects
Community
- LangChain Discord โ Active community discussing RAG and LLMs
- SustainSys AI Academy โ Our other courses on transformers and ML fundamentals
- Research Groups โ Facebook AI, Google Research, DeepMind publish RAG research
Recommended Learning Path
- Complete this course โ
- Read the original RAG paper
- Build a basic RAG system with your own documents
- Study advanced techniques (re-ranking, evaluation)
- Deploy a production system
- Follow research on new RAG techniques