Retrieval-Augmented Generation (RAG)

Retrieval-Augmented Generation (RAG) is a technique that enhances Large Language Models (LLMs) by combining retrieval and generation. Instead of relying solely on a model's training data, RAG systems retrieve relevant documents or knowledge from external sources and use them to augment the generation process, producing more accurate, grounded, and up-to-date responses.

The RAG approach solves critical limitations of standalone LLMs: hallucinations (generating false information), lack of access to real-time or domain-specific knowledge, and inability to trace where information comes from. By integrating a retrieval component, RAG systems can provide citations, handle private documents, and adapt to knowledge updates without retraining.

RAG has become essential in production systems for question-answering, customer support, research assistance, and knowledge work. Companies like OpenAI, Anthropic, and every major AI company have adopted RAG-style approaches in their products. Understanding RAG is critical for anyone building practical AI applications.

What You'll Learn

RAG Architecture

Understand the complete RAG pipeline: document ingestion, chunking, embedding, storage, retrieval, and generation.

Retrieval Systems

Learn vector databases (FAISS, Chroma, Pinecone), embedding models, and retrieval metrics that power modern RAG.

Advanced Techniques

Explore re-ranking, query expansion, multi-hop reasoning, and hybrid search to build production-grade RAG systems.

Real Code & Evaluation

Build RAG systems with LangChain, implement RAGAS evaluation, and learn how to measure RAG quality.

Prerequisites

Familiarity with Python, LLMs (GPT, BERT embeddings), basic machine learning concepts, and ideally some experience with vector databases or information retrieval.

Why RAG Matters

LLMs are powerful but flawed. They hallucinate, have fixed knowledge cutoffs, can't access proprietary documents, and can't update their knowledge without retraining. RAG solves these problems elegantly by separating knowledge retrieval from knowledge generation.

The Problem with LLMs Alone

Hallucinations

LLMs generate plausible-sounding but false information. A 7B parameter model hallucinates ~28% of the time on factual questions.

Stale Knowledge

Models trained in 2023 don't know about 2024 events. Retraining is expensive and impractical for frequent updates.

No Private Data

LLMs can't access your company's documents, databases, or internal knowledge bases without fine-tuning.

No Attribution

Users can't verify where information came from. In regulated industries, this is unacceptable.

How RAG Solves It

RAG elegantly addresses all four problems:

RAG with Citations
96% Verifiable
RAG Factuality
91% Accurate
LLM Only
72% Factual
Real-time Knowledge
100% Current

RAG significantly improves accuracy, attribution, and knowledge currency

RAG in Production

Why companies use RAG: OpenAI uses RAG in ChatGPT's browsing mode. Anthropic uses it for retrieval. Google uses it in Search Generative Experience. Every enterprise AI system uses RAG to ground models in proprietary data and reduce hallucinations.

Historical Evolution of RAG

RAG didn't emerge fully formed in 2020. It represents the convergence of information retrieval research (decades old) with modern deep learning and LLMs.

The Evolution Timeline

1990s-2000s: Information Retrieval Era ▼

BM25 and TF-IDF: Traditional keyword matching dominated. These algorithms ranked documents by term frequency and inverse document frequency. Fast, interpretable, but limited to exact keyword matches.

Vector Space Models: Documents represented as vectors; similarity computed via cosine distance. This foundational idea still underlies modern RAG.

2010-2017: Neural IR and Word Embeddings ▼

Word2Vec (2013): Dense word embeddings showed that semantic relationships could be captured in vector space. This was revolutionary.

Neural Ranking: Researchers began using neural networks to learn ranking functions instead of hand-crafted BM25 weights.

BERT (2018): Contextual embeddings that understand meaning in context, not just individual words.

2018-2020: Attention and QA Systems ▼

BiDAF and Reading Comprehension: Models that could read documents and answer questions about them. This was the precursor to RAG.

Dense Passage Retrieval (DPR, 2020): Facebook AI Research showed that learning dense embeddings specifically for retrieval was more effective than BM25. This was crucial โ€” it showed retrieval itself could be learned end-to-end.

2020-2021: RAG Emerges ▼

RAG Paper (Lewis et al., Facebook, 2020): "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" formally defined RAG. The key insight: condition generation on retrieved documents.

Vector Databases Become Practical: Pinecone (2021), Weaviate, and others made vector databases accessible. No longer required specialized expertise.

LLMs Arrive: GPT-3 (2020) showed that large models could do amazing things. But they still hallucinated.

2022-Present: Modern RAG Era ▼

LangChain (2022): Made RAG accessible to everyone with simple APIs for retrieval, generation, and tool use.

Production RAG at Scale: Companies deployed RAG for customer support, internal search, research, and knowledge work.

Advanced RAG: Re-ranking, query expansion, hypothetical document embeddings, multi-hop reasoning, and more.

RAG Evaluation (RAGAS, 2023): Frameworks for measuring RAG quality. No longer just eyeballing results.

Agentic RAG (2024): Systems that decide when to retrieve, what to retrieve, and iterate on results. RAG becomes a tool in a broader reasoning system.

Core Concepts

To understand RAG, you need to understand a few foundational concepts that come from information retrieval and machine learning.

Embeddings and Semantic Search

An embedding is a numerical representation of text in a high-dimensional space (typically 384-1536 dimensions). Semantically similar texts have embeddings close in vector space.

Cosine Similarity
similarity(A, B) = (A ยท B) / (||A|| ร— ||B||)

Measures angle between vectors. Range: -1 to 1. Higher is more similar.

Embedding Models convert text to vectors:

  • BERT / RoBERTa: Bi-directional context. Slow but accurate. Use for important tasks.
  • Sentence-Transformers: Optimized for semantic search. Fast and effective. Industry standard.
  • OpenAI text-embedding-3-small/large: Proprietary, powerful, but costs add up at scale.
  • Specialized Models: BGE, E5, ColBERT for domain-specific or multi-lingual retrieval.

Vector Databases

Storing and searching embeddings efficiently requires specialized databases optimized for similarity search:

  • FAISS (Facebook): In-memory, blazing fast, used when you fit data in RAM.
  • Chroma: Simple, developer-friendly, good for prototypes and small-to-medium applications.
  • Pinecone: Fully managed, scalable, built for production with metadata filtering.
  • Milvus / Weaviate: Open-source, scalable, support complex queries.
  • Elasticsearch with vector capabilities: For teams that already use Elasticsearch.

The RAG Pipeline: Core Steps

RAG systems follow a standard flow:

RAG Pipeline Flow

๐Ÿ“„
Documents
โ†’
โœ‚๏ธ
Chunk
โ†’
๐Ÿงฎ
Embed
โ†’
๐Ÿ’พ
Store
โ†’
๐Ÿ”
Query
โ†’
๐Ÿ“ฅ
Retrieve
โ†’
โœ๏ธ
Generate

Key Terms

  • Chunking: Breaking documents into overlapping pieces (typically 256-1024 tokens) to manage context length and improve retrieval granularity.
  • Relevance: Whether retrieved documents actually answer the query. Measured by similarity score and human judgment.
  • Ranking: Ordering retrieved results by relevance. Naive RAG ranks by embedding similarity; advanced RAG uses cross-encoders to rerank.
  • Context Window: The maximum tokens the LLM can see. Modern LLMs have 4K-200K tokens; RAG must fit retrieved docs + query + instructions in this budget.
  • Hallucination: LLM generating false information not grounded in retrieved documents. RAG reduces but doesn't eliminate hallucination.

RAG Architecture & Pipeline

A complete RAG system consists of multiple stages. Understanding each stage and how they interact is critical to building effective systems.

The Full RAG Architecture

RAG System Architecture

Document Ingestion PDFs, Web, APIs Databases, Files Preprocessing Chunk & Clean Normalize Embedding Dense Vectors 384-1536 dims Vector Store FAISS, Pinecone Weaviate, Chroma Query Processing Embed Query Same space as docs Retrieval Similarity Search Top-K matching Re-ranking (Optional) Cross-encoder Generation Context + Query LLM generates answer Answer with Citations & Sources Confidence scores

Stage 1: Document Ingestion

Raw documents come from many sources: PDFs, websites, databases, APIs, internal wikis, research papers, product docs. The ingestion system must handle various formats and scale to process millions of documents.

Document Sources

Common sources include customer support docs, research papers, internal knowledge bases, product documentation, website content, and real-time data feeds.

Stage 2: Preprocessing & Chunking

Documents must be broken into manageable pieces (chunks). Too small chunks lose context; too large chunks exceed context windows. Typical chunk size: 256-1024 tokens with 20-50% overlap.

Chunk Overlap Strategy
overlap_pct = 30%
chunk_size = 512
overlap_size = chunk_size ร— overlap_pct
= 154 tokens overlap between chunks

Stage 3: Embedding

Convert text chunks to dense vectors using embedding models. This happens offline once; at query time, queries are embedded in the same space.

Stage 4: Storage & Indexing

Vectors are stored in vector databases with metadata (source document, chunk ID, date, etc.). These databases use specialized indexes (HNSW, IVF, PQ) for fast similarity search.

Stage 5: Query Processing

User query is embedded using the same embedding model. This is critical โ€” using different embedding models for documents and queries breaks the system.

Stage 6: Retrieval

Find top-K similar documents (typically K=3-10) using approximate nearest neighbor search. Trade-off between speed and accuracy.

Stage 7 (Optional): Re-ranking

Use a more powerful model (cross-encoder) to re-rank retrieved results. Slower but more accurate. Optional but recommended for high-quality systems.

Stage 8: Generation

Combine query + retrieved context + system prompt, and feed to LLM. LLM generates answer grounded in retrieved documents.

Key Components of RAG Systems

RAG systems have several critical components. Understanding and optimizing each is key to building production systems.

1. Embedding Models

Sentence-Transformers

Fast, open-source models like all-MiniLM-L6-v2 (384 dims). Great for general-purpose semantic search.

OpenAI Embeddings

Proprietary text-embedding-3 (1536 dims). Powerful but costs increase with scale.

Domain-Specific

BGE for multilingual, ColBERT for specialized domains. Trade speed for accuracy.

Sparse vs Dense

Dense embeddings capture semantics; sparse (BM25) capture exact keywords. Hybrid often works best.

2. Chunking Strategies

Fixed-Size Chunking

Split documents into fixed-size pieces (e.g., 512 tokens). Simple but loses structure. Risk of splitting sentences mid-way.

Semantic Chunking

Use sentence or paragraph boundaries. Preserves meaning but variable size can complicate batch processing.

Recursive Chunking

Split by hierarchy: sentences โ†’ paragraphs โ†’ sections. Preserves semantic structure. Recommended approach.

3. Vector Databases

Choose based on your scale and infrastructure:

Database Scale Speed Cost Best For FAISS ~1M in-memory Ultra-fast Free (compute) Prototypes, single machine Chroma ~100K-1M Fast Free/cheap Development, small apps Pinecone 100M+ Fast (managed) $7-40+/month Production, large scale Weaviate 100M+ Fast Free (self-hosted) Enterprise, open-source Milvus 1B+ Very fast Free (self-hosted) Massive scale, cloud

4. Retrieval Methods

  • Semantic Search: Query and documents in same embedding space. Foundation of RAG.
  • Hybrid Search: Combine dense vectors + sparse BM25. Excellent results, more computation.
  • Metadata Filtering: Retrieve by category, date, or other metadata before similarity search. Reduces hallucination.
  • Multi-hop Retrieval: Iterative retrieval: retrieve โ†’ answer intermediate question โ†’ retrieve again. For complex queries.
  • Query Expansion: Generate multiple versions of query (synonyms, rephrasings) and retrieve for all. Increases recall.

5. Re-ranking (Critical for Quality)

Embedding-based retrieval is fast but can miss nuance. Cross-encoder re-rankers improve relevance:

Re-ranking Pipeline
retrieved_docs = retrieve_top_k(query, k=100)
scored = cross_encoder.predict(
[(query, doc) for doc in retrieved_docs]
)
reranked = sorted_by_score(scored)[:10]

Popular Re-rankers

ms-marco-MiniLM for speed, ms-marco-TinyBERT for extreme speed, mmarco-mMiniLMv2 for multilingual. Each trades accuracy vs latency.

6. Context Management

LLMs have context limits. You must fit: [system prompt] + [retrieved docs] + [query] + [generation space] within limit.

  • Context Pruning: Summarize retrieved docs or keep only most relevant parts.
  • Sliding Window: If too much context, keep top-K most relevant chunks and summarize the rest.
  • Long-Context Models: GPT-4 Turbo (128K), Claude 3 (200K). Higher costs but more flexibility.

Implementation Guide: Building a Basic RAG System

Let's build a working RAG system from scratch. We'll use popular libraries: LangChain for orchestration, Sentence-Transformers for embeddings, and Chroma for vector storage.

Step 1: Install Dependencies

Python โ€” Installation
pip install langchain langchain-community \ sentence-transformers chromadb \ pypdf python-dotenv

Step 2: Load Documents

Python โ€” Document Loading
from langchain_community.document_loaders import PyPDFLoader loader = PyPDFLoader("document.pdf") documents = loader.load() print(f"Loaded {len(documents)} pages") # Documents are split by PyPDFLoader, but we want # more granular chunking from langchain.text_splitter import RecursiveCharacterTextSplitter splitter = RecursiveCharacterTextSplitter( chunk_size=512, # tokens per chunk chunk_overlap=102, # 20% overlap separators=["\n\n", "\n", " ", ""] ) chunks = splitter.split_documents(documents) print(f"Created {len(chunks)} chunks")

Step 3: Create Embeddings and Store in Vector DB

Python โ€” Embedding and Storage
from sentence_transformers import SentenceTransformer from langchain_community.vectorstores import Chroma # Initialize embedding model embedding_model = SentenceTransformer( "all-MiniLM-L6-v2" ) # Create vector store vectorstore = Chroma.from_documents( documents=chunks, embedding=embedding_model, collection_name="documents", persist_directory="./chroma_db" ) print(f"Stored {len(chunks)} chunks in vector database") vectorstore.persist() # Save to disk

Step 4: Create a Retriever

Python โ€” Retriever Setup
# Create retriever with MMR (Maximal Marginal Relevance) # This balances relevance with diversity retriever = vectorstore.as_retriever( search_type="mmr", search_kwargs={ "k": 5, # Return top 5 results "fetch_k": 20, # Fetch 20 then select 5 } ) # Test retrieval query = "What is machine learning?" retrieved = retriever.get_relevant_documents(query) print(f"Retrieved {len(retrieved)} documents") for doc in retrieved: print(f"- {doc.page_content[:100]}...")

Step 5: Set Up LLM and Build RAG Chain

Python โ€” RAG Chain
from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.schema.runnable import RunnablePassthrough from langchain.schema.output_parser import StrOutputParser # Initialize LLM llm = ChatOpenAI( model="gpt-3.5-turbo", temperature=0.2, api_key="your-api-key" ) # Create prompt template prompt_template = """Use the following documents to answer the question. If you can't find the answer in the documents, say "I don't know." Documents: {context} Question: {question} Answer:""" prompt = PromptTemplate( template=prompt_template, input_variables=["context", "question"] ) # Create RAG chain def format_docs(docs): return "\n\n".join([d.page_content for d in docs]) chain = ( {"context": retriever | format_docs, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() ) # Query! result = chain.invoke("What is machine learning?") print(result)

Step 6: Add Source Attribution

Python โ€” Source Attribution
# Modify chain to return docs and answer from langchain.schema.runnable import RunnableMap def create_rag_with_sources(retriever, llm, prompt): return RunnableMap({ "docs": retriever, "question": RunnablePassthrough() }) | { "answer": ( {"context": lambda x: format_docs(x["docs"]), "question": lambda x: x["question"]} | prompt | llm | StrOutputParser() ), "sources": lambda x: x["docs"] } # Use it rag_with_sources = create_rag_with_sources(retriever, llm, prompt) result = rag_with_sources.invoke("What is machine learning?") print("Answer:", result["answer"]) print("\nSources:") for doc in result["sources"]: print(f"- {doc.metadata.get('source', 'Unknown')}")

Advanced RAG Techniques

Basic RAG is powerful, but production systems need advanced techniques to achieve high quality. Here are the most important optimizations.

1. Re-ranking with Cross-Encoders

Initial retrieval via embeddings is fast but imprecise. Cross-encoders compute interaction scores between query and document, providing more accurate relevance scores.

Python โ€” Cross-Encoder Re-ranking
from sentence_transformers import CrossEncoder # Load cross-encoder cross_encoder = CrossEncoder( "ms-marco-MiniLM-L-6-v2" ) # Retrieve many documents retrieved = retriever.get_relevant_documents(query) # Score with cross-encoder pairs = [(query, doc.page_content) for doc in retrieved] scores = cross_encoder.predict(pairs) # Sort by cross-encoder scores ranked_docs = sorted( zip(retrieved, scores), key=lambda x: x[1], reverse=True )[:5] # Keep top 5 print("Re-ranked results:") for doc, score in ranked_docs: print(f"Score: {score:.3f} - {doc.page_content[:80]}...")

2. Query Expansion

Generate multiple versions of the query to improve recall. Especially useful for complex or ambiguous questions.

Python โ€” Query Expansion
from langchain.prompts import PromptTemplate from langchain.llms import OpenAI # Create query expander expansion_template = """Generate 3 alternative phrasings of the following question that would help find relevant documents. Return only the phrasings, one per line. Original question: {question} Alternatives:""" expander_prompt = PromptTemplate( template=expansion_template, input_variables=["question"] ) # Expand query llm = OpenAI(temperature=0.7) expanded = (expander_prompt | llm).invoke(question=original_query) # Split into individual queries queries = [original_query] + [ q.strip() for q in expanded.split("\n") if q.strip() ] # Retrieve for all queries and combine all_docs = [] for q in queries: docs = retriever.get_relevant_documents(q) all_docs.extend(docs) # Remove duplicates while preserving order seen = set() unique_docs = [] for doc in all_docs: doc_id = doc.metadata.get("source") + doc.page_content[:50] if doc_id not in seen: seen.add(doc_id) unique_docs.append(doc) print(f"Retrieved {len(unique_docs)} unique documents from {len(queries)} queries")

3. Hypothetical Document Embeddings (HyDE)

Instead of embedding the query directly, generate a hypothetical answer, embed that, then use it to retrieve. This works surprisingly well!

Python โ€” HyDE Implementation
# Generate hypothetical document hyde_prompt = """Write a short document that would answer the following question: Question: {question} Document:""" hyde_template = PromptTemplate( template=hyde_prompt, input_variables=["question"] ) # Generate hypothetical doc llm = OpenAI(temperature=0.7) hypothetical_doc = (hyde_template | llm).invoke(question=query) # Embed the hypothetical document from sentence_transformers import SentenceTransformer embedder = SentenceTransformer("all-MiniLM-L6-v2") query_embedding = embedder.encode(hypothetical_doc) # Retrieve using this embedding retrieved = vectorstore.similarity_search_by_vector( query_embedding, k=5 ) print("HyDE-based retrieval:") for doc in retrieved: print(f"- {doc.page_content[:100]}...")

4. Multi-hop Retrieval (Iterative RAG)

For complex questions requiring multiple steps, retrieve โ†’ answer intermediate question โ†’ retrieve again.

Python โ€” Multi-hop Retrieval
def multi_hop_retrieval(question, retriever, llm, max_hops=3): """Iteratively retrieve and reason about intermediate questions.""" context = [] current_question = question for hop in range(max_hops): # Retrieve documents for current question docs = retriever.get_relevant_documents(current_question) context.extend(docs) # Check if we have enough info doc_text = "\n\n".join([d.page_content for d in docs]) check_prompt = f"""Based on these documents: {doc_text} Can you fully answer: {question}? Respond with: 1. YES if fully answered 2. NO if more info needed 3. If NO, what should we search for next?""" response = llm(check_prompt) if "YES" in response.upper(): break elif "NO" in response.upper(): # Extract next search query next_query_prompt = f"""Extract one search query to find missing information. Be specific. {response} Search query:""" current_question = llm(next_query_prompt) else: break return context # Use it context = multi_hop_retrieval( "How does climate change affect agriculture in Africa?", retriever, llm ) print(f"Retrieved {len(context)} documents across multiple hops")

5. Hybrid Search (Dense + Sparse)

Combine semantic search (dense vectors) with keyword search (BM25). Hybrid often outperforms either alone.

Python โ€” Hybrid Search
from langchain_community.retrievers import BM25Retriever from langchain.retrievers import EnsembleRetriever # Create dense retriever (semantic) dense_retriever = vectorstore.as_retriever( search_kwargs={"k": 10} ) # Create sparse retriever (BM25 keyword search) bm25_retriever = BM25Retriever.from_documents(chunks) bm25_retriever.k = 10 # Combine them hybrid_retriever = EnsembleRetriever( retrievers=[dense_retriever, bm25_retriever], weights=[0.5, 0.5] # Adjust weights based on your data ) # Use hybrid retriever docs = hybrid_retriever.get_relevant_documents(query) print(f"Retrieved {len(docs)} documents (hybrid search)")

6. Metadata Filtering

Filter retrieval by metadata (source, date, category) before similarity search. Reduces hallucinations.

Python โ€” Metadata Filtering
# Retrieve with metadata filters docs = vectorstore.similarity_search( "machine learning", k=5, filter={"source": "research_papers"} # Only from research papers ) # More complex filters docs = vectorstore.similarity_search( "recent advances", k=5, filter={ "$and": [ {"year": {"$gte": 2023}}, {"category": {"$in": ["AI", "ML"]}} ] } ) print(f"Retrieved {len(docs)} documents with metadata filters")

Comparing RAG Approaches

There are multiple ways to implement RAG. Each has tradeoffs in complexity, quality, latency, and cost.

Naive RAG vs Advanced RAG vs Agentic RAG

Aspect Naive RAG Advanced RAG Agentic RAG Retrieval Single query embedding Query expansion + re-ranking Iterative decisions on what/when to retrieve Architecture Query โ†’ Retrieve โ†’ Generate Query โ†’ Expand โ†’ Retrieve โ†’ Rerank โ†’ Generate Reasoning loop with retrieval as a tool Latency 100-200ms (fastest) 500-2000ms 2-10s (depends on reasoning steps) Quality ~70-75% factuality ~85-92% factuality ~90-95% (best) Cost $0.001-0.01 per query $0.01-0.05 per query $0.05-0.50 per query (multiple LLM calls) Complexity Low (few components) Medium (orchestration needed) High (reasoning, tool use) Use Cases Quick prototypes, simple QA Production systems, enterprise search Complex reasoning, multi-step tasks

When to Use Each

Naive RAG

Use for MVP/prototypes, real-time applications where latency is critical, or simple fact retrieval. Good for: chatbots answering FAQs, simple document search.

Advanced RAG

Use for production systems where quality matters, enterprise knowledge bases, research assistants. Good for: customer support, internal search, documentation QA.

Agentic RAG

Use for complex reasoning, multi-step problem solving, when decisions about what/when to retrieve are important. Good for: research assistance, complex analysis, reasoning-heavy tasks.

Real-World RAG Use Cases

RAG is transforming how organizations interact with information. Here are the most impactful applications.

Customer Support & Q&A

Support Ticket Automation

RAG searches support docs, FAQs, and past tickets to suggest or auto-generate responses. Reduces response time by 50-80%, reduces support costs significantly.

Interactive FAQ

Users ask questions in natural language; RAG retrieves relevant FAQ entries and generates comprehensive answers with citations.

Chatbot Personalization

RAG retrieves customer history, account details, and relevant docs to provide personalized support. Higher satisfaction scores.

Enterprise Knowledge Management

Internal Search

Employees search across documentation, wikis, and policies. RAG understands intent and returns relevant information with context.

Research Assistant

Scientists and analysts use RAG to search papers, data, and findings. Helps synthesize information across thousands of documents.

Policy Compliance

RAG helps identify relevant policies and regulations for queries. Important for legal, finance, and regulated industries.

Content & Creative

Content Generation

RAG retrieves source material (research, interviews, data) to ground content generation. Reduces fabrication, improves accuracy.

Code Search

Developers search codebases using natural language. RAG finds relevant code snippets and documentation. GitHub Copilot uses this.

Product Recommendations

RAG retrieves product info, reviews, and specs. Generates personalized recommendations with explanations.

Healthcare & Regulatory

Clinical Decision Support

RAG retrieves medical literature, treatment guidelines, and patient history. Assists doctors in decision making. Requires high accuracy.

Drug Discovery

Researchers use RAG to search papers on compounds, mechanisms, and trials. Accelerates identification of research directions.

Regulatory Compliance

Retrieve relevant regulations, requirements, and precedents. Critical for healthcare, finance, and pharma.

E-commerce & Marketplace

Smart Product Search

RAG understands shopper intent and retrieves relevant products from detailed catalogs. Improves conversion by 15-30%.

Seller Support

Marketplace sellers use RAG to search policies, shipping rules, and best practices. Reduces policy violations.

Enterprise RAG Applications

Enterprise RAG systems must handle scale, security, compliance, and integration with existing systems. This section covers enterprise-grade patterns.

Scalability Patterns

Multi-Collection Architecture

Store different data sources in separate vector collections. Allows different embedding models, refresh rates, and metadata schemas. Query relevant collections based on intent.

Sharding

For massive document collections (100M+), shard across multiple vector databases. Route queries to relevant shards. Requires careful metadata strategy.

Hybrid Deployment

Mix cloud-based vector DBs (Pinecone) with on-premises or private deployments for sensitive data. Use federated search.

Security & Privacy

  • Data Isolation: Separate collections per tenant/organization. Ensure queries only retrieve authorized documents.
  • Encryption: Encrypt documents at rest and in transit. Vector DBs support encrypted storage.
  • Access Control: Implement row-level security. Tag documents with access levels and filter during retrieval.
  • Audit Logging: Log all queries and retrievals for compliance and debugging.
  • PII Handling: Anonymize or mask sensitive data in retrieved documents before sending to LLM.

Quality Assurance

Evaluation Metrics

Track retrieval accuracy (recall, precision), generation quality (factuality, relevance), and end-to-end user satisfaction.

Feedback Loops

Collect user feedback on answer quality. Use feedback to identify failing queries and improve chunking/retrieval strategies.

A/B Testing

Test different embeddings, chunk sizes, retrieval strategies, and LLMs. Measure impact on quality and latency.

Integration with Enterprise Systems

Data Connectors

Integrate with Salesforce, Jira, Confluence, SharePoint, databases. Keep vector store in sync with source data.

API Gateway

Expose RAG system via API. Implement rate limiting, authentication, and request logging.

Workflow Integration

Embed RAG into existing workflows. Trigger retrieval from document management systems, CRM, or chat platforms.

Cost Management

  • Cache Retrieved Documents: Store frequently retrieved docs in cache. Reduces redundant embedding computations.
  • Batch Embedding: Embed documents in batches during off-peak hours, not in real-time.
  • Model Selection: Use smaller, faster embedding models (384-dim) instead of large ones when possible. Fine-tune on domain data for better quality without scaling to larger models.
  • Selective Re-ranking: Only re-rank top-20 results, not all results.
  • Query Optimization: Cache embeddings of common queries. Deduplicate similar queries before retrieving.

Common RAG Mistakes (And How to Avoid Them)

Building RAG systems is deceptively simple but deploying high-quality production systems is hard. Here are the most common pitfalls.

1. Using Different Embedding Models for Indexing vs Querying

Mistake: Index with OpenAI embeddings but query with Sentence-Transformers. Why Bad: Vectors are in different spaces. Similarity scores are meaningless. Fix: Use the same model for indexing and querying, always.

2. Ignoring Chunk Overlap and Context

Mistake: Split documents into 256-token chunks with 0% overlap. Why Bad: Important concepts get split across chunks, retrieved documents lack context. Fix: Use 20-50% overlap. Test different chunk sizes.

3. Not Evaluating Retrieval Quality

Mistake: Only evaluate end-to-end answer quality. Why Bad: Can't diagnose if problems are in retrieval or generation. Fix: Evaluate retrieval metrics (recall, precision, NDCG) separately from generation quality (factuality, relevance).

4. Over-Optimizing for Latency

Mistake: Use small embeddings, no re-ranking, minimal retrieval to reduce latency. Why Bad: Answer quality suffers. Fix: Measure quality-latency tradeoff. Users prefer accurate over fast.

5. Forgetting About LLM Hallucinations Even With RAG

Mistake: Assume retrieved documents prevent hallucinations. Why Bad: LLMs still hallucinate even when given correct context. Happen ~10-20% of the time even with good retrieval. Fix: Implement fact-checking, confidence scoring, and require citations.

6. Not Handling Updates to Documents

Mistake: Embed documents once and never update. Why Bad: System becomes stale. Critical for real-time or frequently updated data. Fix: Implement incremental updates. For some data, re-embed daily or hourly.

7. Using Raw Documents Without Metadata

Mistake: Store chunks with no metadata (source, date, category, access level). Why Bad: Can't filter, can't trace results, violates compliance. Fix: Rich metadata on every chunk. Use for filtering and attribution.

8. Ignoring the Cost of Embeddings at Scale

Mistake: Use expensive embedding APIs without considering cost. Why Bad: At 1M documents, cost becomes prohibitive. Fix: Use open-source models (Sentence-Transformers). Only use proprietary APIs when necessary.

9. Not Testing on Your Actual Data

Mistake: Test RAG on toy datasets, deploy on real data. Why Bad: Real data has different characteristics. Quality may be much lower. Fix: Evaluate on representative samples of your actual data early and often.

10. Underestimating the Importance of Query Understanding

Mistake: Embed query as-is without preprocessing. Why Bad: Typos, abbreviations, and poor phrasing hurt retrieval. Fix: Implement query expansion, spelling correction, and intent classification before retrieval.

RAG Best Practices

Based on what works in production, here are proven best practices for building high-quality RAG systems.

Document Preparation

  • Clean Data: Remove headers, footers, boilerplate. Preserve semantic structure. Quality in โ†’ quality out.
  • Rich Metadata: Attach source, date, category, confidence, access level, version. Use for filtering and attribution.
  • Optimal Chunking: 256-1024 tokens is typical. Test different sizes. Recursive chunking > fixed-size chunking.
  • Overlap: 20-50% overlap between chunks. Helps maintain context at chunk boundaries.

Embedding Selection

  • Start with Sentence-Transformers: all-MiniLM-L6-v2 (384-dim) is a strong baseline. Free, fast, effective.
  • Fine-tune if needed: If domain-specific vocabulary is important, fine-tune an embedding model on your data.
  • Multilingual: If serving multiple languages, use multilingual models (multilingual-e5-small).
  • Consistency: Same model for indexing and querying. Document which model was used.

Retrieval Strategy

  • Start Simple: Semantic search (dense retrieval) works well. Add complexity only if needed.
  • Hybrid Search for Recall: Combine dense + sparse (BM25). Higher recall with minimal latency increase.
  • Always Re-rank: If quality is critical, re-rank top-20 with cross-encoder. ~500ms extra latency, 10-20% quality improvement.
  • Metadata Filters: Filter by date, source, category before similarity search. Reduces noise and hallucination.

Generation

  • Context Window Management: Fit [prompt + retrieved docs + generation] in context. Prioritize retrieved docs, summarize if needed.
  • System Prompt: Clear instructions to use retrieved docs, cite sources, say "I don't know" if not in docs.
  • Temperature: Low temperature (0.2-0.3) for factual tasks. Higher (0.7+) for creative tasks.
  • Stopping Tokens: Use stop tokens if generating citations in specific format.

Evaluation & Monitoring

  • Retrieval Metrics: Recall@K, Precision@K, NDCG. Essential for diagnosing issues.
  • Generation Metrics: Use RAGAS (Retrieval-Augmented Generation Assessment) framework. Measures answer relevance, groundedness, and citation accuracy.
  • User Feedback: Thumbs up/down on answers. Use to identify and fix failing queries.
  • Continuous Testing: Build evaluation pipeline. Test weekly on new data.

Deployment

  • Start Small: Begin with small dataset, expand after validating quality.
  • Version Control: Version embeddings and chunking strategy. A/B test new versions.
  • Monitoring: Track latency, error rates, cost. Alert on quality degradation.
  • User Experience: Show confidence scores, source documents, alternative answers.

Advanced Insights & Future Directions

RAG is rapidly evolving. Here are emerging techniques and future directions shaping the field.

Sparse-Dense Hybrids

Recent research shows that hybrid retrieval (combining sparse BM25 with dense vectors) often beats either alone. More nuanced: different retrieval methods excel at different aspects (density captures semantics, sparsity captures keywords). Future systems will intelligently choose methods based on query type.

Self-Improving RAG

Systems that learn from retrieval failures. When answer is wrong, identify whether fault was in retrieval or generation. Adapt: better query expansion for retrieval failures, better prompts for generation failures.

Multimodal RAG

Extending RAG beyond text: retrieve images, tables, code snippets, videos. Require multimodal embeddings (CLIP, LLaVA) and multimodal generation models. Frontier is still being mapped.

Graph-based RAG

Instead of flat document retrieval, build knowledge graphs from documents. Retrieve subgraphs relevant to query. Allows reasoning across multiple documents and relationships. More complex but higher quality for knowledge-intensive tasks.

Long-Context Models Changing the Game

GPT-4 Turbo and Claude 3 have 100K-200K context windows. Some researchers are experimenting with putting entire corpora in context (in-context RAG). Extreme approach: index everything in the prompt. Reduces retrieval quality concerns but increases cost and latency.

Agentic RAG

RAG as a tool in reasoning systems. LLM decides what to retrieve, when, and how many times. Learns to reason about information gaps. Early results show promise but current LLMs struggle with this level of metacognition.

Fine-Tuning on Retrieval Tasks

Instruction-tuned models specifically for retrieval-augmented generation tasks. E.g., training models to generate better queries, to recognize when they need to retrieve, or to synthesize information from multiple docs.

Latency Optimization

Current bottleneck: embedding the query and retrieving from vector DB (100-200ms). Future: cached embeddings, streaming retrieval, early stopping. Goal: sub-50ms retrieval for real-time applications.

Complete Code Examples

Working, production-ready code for building RAG systems. Copy-paste and run.

Example 1: Basic RAG with LangChain

Python โ€” Basic RAG Pipeline
import os from langchain.document_loaders import PyPDFLoader from langchain.text_splitter import RecursiveCharacterTextSplitter from langchain.embeddings import HuggingFaceEmbeddings from langchain.vectorstores import Chroma from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.schema.runnable import RunnablePassthrough from langchain.schema.output_parser import StrOutputParser # Set API key os.environ["OPENAI_API_KEY"] = "your-key-here" # 1. Load documents loader = PyPDFLoader("example.pdf") documents = loader.load() # 2. Chunk documents splitter = RecursiveCharacterTextSplitter( chunk_size=512, chunk_overlap=102 ) chunks = splitter.split_documents(documents) # 3. Create embeddings and vector store embeddings = HuggingFaceEmbeddings( model_name="all-MiniLM-L6-v2" ) vectorstore = Chroma.from_documents( chunks, embeddings, collection_name="rag-docs", persist_directory="./chroma_data" ) # 4. Create retriever retriever = vectorstore.as_retriever( search_type="mmr", search_kwargs={"k": 5, "fetch_k": 20} ) # 5. Setup LLM and RAG chain llm = ChatOpenAI(model="gpt-3.5-turbo", temperature=0.2) prompt_template = """Use the documents to answer the question. If you can't answer from the documents, say "I don't know." Documents: {context} Question: {question} Answer:""" prompt = PromptTemplate( template=prompt_template, input_variables=["context", "question"] ) def format_docs(docs): return "\n\n".join([d.page_content for d in docs]) chain = ( {"context": retriever | format_docs, "question": RunnablePassthrough()} | prompt | llm | StrOutputParser() ) # 6. Query result = chain.invoke("What is the main topic of the document?") print(result)

Example 2: RAG with Re-ranking

Python โ€” RAG with Cross-Encoder Re-ranking
from langchain.vectorstores import Chroma from sentence_transformers import CrossEncoder from langchain.chat_models import ChatOpenAI # Setup retriever and re-ranker vectorstore = Chroma( collection_name="rag-docs", embedding_function=embeddings, persist_directory="./chroma_data" ) retriever = vectorstore.as_retriever(search_kwargs={"k": 20}) cross_encoder = CrossEncoder("ms-marco-MiniLM-L-6-v2") # Setup LLM llm = ChatOpenAI(model="gpt-3.5-turbo") def retrieve_and_rerank(query): # Initial retrieval docs = retriever.get_relevant_documents(query) # Re-rank with cross-encoder pairs = [(query, doc.page_content) for doc in docs] scores = cross_encoder.predict(pairs) # Sort and keep top-5 scored_docs = sorted( zip(docs, scores), key=lambda x: x[1], reverse=True )[:5] return [doc for doc, score in scored_docs] # Use in chain docs = retrieve_and_rerank("your query here") context = "\n\n".join([d.page_content for d in docs]) print(f"Retrieved {len(docs)} documents after re-ranking")

Example 3: Advanced RAG with Query Expansion and Filtering

Python โ€” Advanced RAG with Query Expansion
from langchain.chat_models import ChatOpenAI from langchain.prompts import PromptTemplate from langchain.output_parsers import CommaSeparatedListOutputParser # Query expansion expansion_template = """Generate 3 alternative phrasings of this query: {original_query} Alternatives:""" expansion_prompt = PromptTemplate( template=expansion_template, input_variables=["original_query"] ) llm = ChatOpenAI(temperature=0.7) parser = CommaSeparatedListOutputParser() # Expand query original_query = "How does RAG improve LLMs?" expanded = (expansion_prompt | llm | parser).invoke( original_query=original_query ) # Retrieve with all queries all_queries = [original_query] + expanded retrieved_docs = [] seen = set() for q in all_queries: docs = retriever.get_relevant_documents(q) for doc in docs: doc_id = hash(doc.page_content) if doc_id not in seen: seen.add(doc_id) retrieved_docs.append(doc) print(f"Query expansion: {len(all_queries)} queries โ†’ {len(retrieved_docs)} unique docs")

Example 4: RAG Evaluation with RAGAS

Python โ€” RAGAS Evaluation Framework
pip install ragas from ragas import evaluate from ragas.metrics import ( faithfulness, answer_relevancy, context_relevancy, context_precision ) # Prepare evaluation data eval_data = { "question": [ "What is RAG?", "How does retrieval improve LLMs?" ], "answer": [ "RAG combines retrieval and generation...", "Retrieval provides grounding..." ], "contexts": [ ["RAG is a technique...", "It uses vector..."], ["Retrieval augments...", "Reduces hallucination..."] ], "ground_truth": [ "RAG is Retrieval-Augmented Generation", "It prevents hallucinations" ] } # Evaluate results = evaluate( eval_data, metrics=[ faithfulness, answer_relevancy, context_relevancy, context_precision ] ) print(f"Faithfulness: {results['faithfulness']:.3f}") print(f"Answer Relevancy: {results['answer_relevancy']:.3f}") print(f"Context Relevancy: {results['context_relevancy']:.3f}")

Hands-On Exercises

Learning RAG is best done by building. Here are practical exercises to deepen your understanding.

Exercise 1: Build Your First RAG System

Create a Basic RAG Pipeline

Load a PDF document (download any research paper from arXiv), chunk it, embed it, and build a RAG system that answers questions about it. Steps: 1. Find a research paper PDF (aim for 5-15 pages) 2. Use PyPDFLoader to load it 3. Chunk with RecursiveCharacterTextSplitter (512 tokens, 20% overlap) 4. Create embeddings with Sentence-Transformers 5. Store in Chroma 6. Build QA chain with LangChain 7. Test with 5-10 questions about the paper Success Criteria: System answers at least 80% of questions correctly with proper citations.

Python โ€” Starter Code
# Template loader = PyPDFLoader("paper.pdf") docs = loader.load() splitter = RecursiveCharacterTextSplitter(...) chunks = splitter.split_documents(docs) # Continue from here...

Exercise 2: Compare Embedding Models

Evaluate Different Embeddings

Take the same documents and create vector stores with 3 different embedding models. Compare retrieval quality on the same test queries. Models to compare: - all-MiniLM-L6-v2 (fast, small) - all-mpnet-base-v2 (balanced) - all-distilroberta-v1 (small, decent) Metrics: - Retrieval latency - Relevance (manual scoring 1-5) - Memory usage What you'll learn: How to choose embeddings for different requirements.

Exercise 3: Advanced Retrieval with Re-ranking

Implement Cross-Encoder Re-ranking

Add a re-ranking layer to your RAG system. Compare quality before/after re-ranking. Setup: 1. Use your basic RAG from Exercise 1 2. Retrieve top-20 documents (instead of top-5) 3. Re-rank with ms-marco-MiniLM cross-encoder 4. Keep top-5 re-ranked results Evaluation: - Does re-ranking improve answer quality? - How much latency cost? - Is it worth it for your use case?

Exercise 4: Implement Query Expansion

Build Query Expansion

Implement query expansion: given a query, generate 3 alternative phrasings and retrieve from all. Compare to single-query retrieval. Process: 1. Use an LLM to expand queries 2. Retrieve for original + expanded 3. Deduplicate results Measure: - Recall improvement - Latency increase

Exercise 5: Evaluate with RAGAS

Setup RAG Evaluation

Install RAGAS and evaluate your RAG system on 20 test questions. Metrics to track: - Faithfulness (does answer follow context?) - Answer relevancy (does it answer the question?) - Context precision (are retrieved docs relevant?) Goal: Achieve >0.85 on all metrics.

RAG Interview Questions

Preparing for interviews? Here are the questions you're likely to face on RAG systems.

Foundational Questions

1. What is RAG and why is it important? ▼
RAG (Retrieval-Augmented Generation) combines document retrieval with LLM generation. It's important because it addresses LLM hallucinations, adds domain-specific knowledge, and provides attribution. Without RAG, LLMs are limited to training data and tend to hallucinate.
2. How does RAG differ from fine-tuning? ▼
Fine-tuning updates model weights on task-specific data (expensive, slow, permanent). RAG retrieves at inference time (flexible, fast to update, cheaper). RAG is better for frequently-updated knowledge; fine-tuning for task adaptation.
3. What are embeddings and why are they central to RAG? ▼
Embeddings are dense vector representations of text in high-dimensional space (typically 384-1536 dimensions). Semantically similar texts have similar embeddings. RAG relies on embeddings to find relevant documents via similarity search.

Architecture & Design Questions

4. Walk me through a complete RAG pipeline. What happens at each stage? ▼
Ingestion: Load documents. Chunking: Split into 256-1024 token pieces with overlap. Embedding: Convert to vectors. Storage: Store in vector DB. Query: Embed user query. Retrieval: Find top-K similar docs. Generation: Feed docs + query to LLM.
5. Why can't we just put all documents in the LLM context? ▼
LLMs have finite context windows (4K-200K tokens). Putting all documents is infeasible for large corpora, increases latency and cost, and actually degrades quality (longer context โ†’ more noise). Retrieval solves this by selecting only relevant docs.
6. What's the difference between embedding a document vs embedding a query? Can you use the same model? ▼
You must use the same model for both. If you use different models, documents and queries are in different vector spaces, making similarity scores meaningless. This is a common mistake.

Technical Implementation

7. How would you chunk documents? What trade-offs are there? ▼
Chunk size 256-1024 tokens typical. Small chunks: preserve relevance, fit more in context, but may lose context. Large chunks: preserve context, retrieve fewer docs, but may exceed context window. Use recursive chunking (by hierarchy) rather than fixed-size. 20-50% overlap helps with boundary cases.
8. How do you choose an embedding model? ▼
Start with all-MiniLM-L6-v2 (fast, 384-dim, effective). Consider: latency, quality, cost, multilingual needs, domain. Fine-tune on domain data if needed. Use proprietary APIs only when necessary due to cost.
9. What vector database would you choose? Why? ▼
Depends on scale: FAISS for <1M in-memory, Chroma for <1M on disk, Pinecone for 100M+. Milvus/Weaviate for open-source at scale. For most startups: Chroma initially, Pinecone when scaling.

Quality & Advanced Topics

10. How do you improve RAG quality? ▼
Retrieval: re-ranking, query expansion, hybrid search. Generation: better prompts, few-shot examples, instructions to cite. Evaluation: track retrieval and generation metrics separately. A/B test changes.
11. What's the difference between naive RAG and advanced RAG? ▼
Naive RAG: single query embedding โ†’ retrieve. Advanced RAG: query expansion โ†’ hybrid search โ†’ re-ranking โ†’ generate. Higher quality but more latency and complexity.
12. How do you handle hallucinations in RAG? ▼
RAG reduces but doesn't eliminate hallucinations. Strategies: clear system prompts telling model to use retrieved docs only, confidence scoring, fact-checking, requiring citations, user feedback loops. Monitor hallucination rate.
13. How would you make RAG work with proprietary/sensitive documents? ▼
Use on-premises vector DB or secure cloud storage. Encrypt at rest and in transit. Implement row-level security โ€” filter retrieved docs by user permissions. Use private LLMs if data is very sensitive. Audit logging.
14. What metrics would you track for a production RAG system? ▼
Retrieval: recall@K, precision@K. Generation: faithfulness, relevancy, hallucination rate. End-to-end: user satisfaction, cost per query, latency. Setup dashboards and alerts.

Frequently Asked Questions

How much data do I need to build a RAG system? ▼
None! RAG works with any amount of data, from a few documents to millions. Start small (10-100 docs) to test ideas. Quality of documents matters more than quantity.
Will RAG make my LLM never hallucinate? ▼
No. RAG reduces hallucinations from ~30% to ~5-10%. But LLMs can still hallucinate even with retrieved context. Implement checks: fact verification, confidence scoring, requiring citations.
How long does it take to build a production RAG system? ▼
MVP (basic RAG): 1-2 weeks. Production-grade (re-ranking, evaluation, monitoring): 2-3 months. Enterprise (security, compliance, integration): 3-6 months.
Can I use RAG with open-source LLMs? ▼
Yes! RAG is architecture-agnostic. Works with GPT, Claude, Llama, Mistral, etc. Open-source LLMs: Llama 2, Mistral, are good choices for cost-sensitive deployments.
How do I handle documents longer than the context window? ▼
Chunk into smaller pieces. Retrieve only relevant chunks. Summarize if needed. Or use newer long-context models (Claude 3 with 200K tokens).
What if my data changes frequently? ▼
Update vector DB incrementally. Add new docs, update changed docs, delete removed docs. Some systems re-embed weekly. For real-time data, consider streaming pipelines.
How do I handle multiple languages? ▼
Use multilingual embedding models (multilingual-e5-base, mmarco-mMiniLMv2). They embed documents and queries in shared space across languages. LLM must support target languages.
Can RAG handle images, tables, code? ▼
Text: yes. Multimodal (images, tables): emerging. Use OCR for images/tables โ†’ text, then standard RAG. For specialized needs: multimodal embeddings (CLIP), multimodal LLMs (GPT-4V).
How much does RAG cost at scale? ▼
Vector DB: $100-10k/month depending on volume. Embeddings: $10-100/month (open-source free). LLM API: $1-100k+/month depending on queries. Total: often cheaper than fine-tuning or large models.
What's the latency of RAG systems? ▼
Typical: 500-2000ms. Breakdown: retrieval (100-300ms), re-ranking (300-500ms), LLM generation (200-1000ms). Can optimize to <500ms for simple systems.

Summary

RAG is the practical solution for building AI systems that are accurate, grounded, and up-to-date. Here's what you've learned:

Key Takeaways

RAG Fundamentals

RAG combines retrieval (finding relevant docs) and generation (LLM creating answers). Solves hallucination, stale knowledge, and attribution problems.

Complete Pipeline

Document ingestion โ†’ chunking โ†’ embedding โ†’ storage โ†’ query processing โ†’ retrieval โ†’ generation. Each stage matters.

Production Patterns

Naive RAG is simple but limited. Advanced RAG (re-ranking, expansion) gives 15-20% quality boost. Choose based on requirements.

Evaluation

Measure retrieval quality (recall, precision) separately from generation (faithfulness, relevancy). Use RAGAS framework.

Next Steps

  1. Build a basic RAG system using the code examples in this guide.
  2. Evaluate on your data using RAGAS or similar frameworks.
  3. Optimize retrieval with re-ranking and query expansion.
  4. Deploy to production with monitoring and user feedback loops.
  5. Stay updated โ€” RAG is evolving rapidly with new techniques emerging monthly.

Common Pitfalls to Avoid

  • Different embedding models for indexing vs querying
  • Not evaluating retrieval quality separately
  • Expecting RAG to completely eliminate hallucinations
  • Ignoring updates to source documents
  • Not tracking costs at scale

Resources for Continued Learning

  • Papers: "Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks" (Lewis et al.) โ€” the seminal RAG paper
  • Libraries: LangChain, LlamaIndex for RAG orchestration
  • Evaluation: RAGAS framework for systematic evaluation
  • Communities: LangChain Discord, RAG research papers on arXiv

Resources & Further Learning

Essential Papers

Tools & Libraries

  • LangChain (langchain.com) โ€” Python library for RAG orchestration. Most popular.
  • LlamaIndex (llamaindex.ai) โ€” Alternative RAG framework, strong for document indexing.
  • Sentence-Transformers (sbert.net) โ€” Open-source embedding models.
  • Chroma (trychroma.com) โ€” Vector database for development.
  • Pinecone (pinecone.io) โ€” Managed vector database for production.
  • RAGAS (github.com/explodinggradients/ragas) โ€” Evaluation framework for RAG systems.
  • DSPy โ€” Formal framework for composing RAG programs.

Learning Resources

  • LangChain Documentation โ€” Best practices and tutorials
  • Hugging Face Hub โ€” Thousands of pre-trained embedding models
  • arXiv.org โ€” Latest research papers on retrieval and generation
  • GitHub โ€” See real implementations in production projects

Community

  • LangChain Discord โ€” Active community discussing RAG and LLMs
  • SustainSys AI Academy โ€” Our other courses on transformers and ML fundamentals
  • Research Groups โ€” Facebook AI, Google Research, DeepMind publish RAG research

Recommended Learning Path

  1. Complete this course โœ“
  2. Read the original RAG paper
  3. Build a basic RAG system with your own documents
  4. Study advanced techniques (re-ranking, evaluation)
  5. Deploy a production system
  6. Follow research on new RAG techniques