Chunking Strategies for RAG Systems
Chunking is the foundational step in building effective Retrieval-Augmented Generation (RAG) systems. It's the process of intelligently splitting large documents into smaller, semantically coherent pieces that can be embedded, indexed, and retrieved to answer user queries. Done well, chunking dramatically improves retrieval accuracy and answer quality. Done poorly, it degrades both dramatically.
The naive approach โ splitting documents on arbitrary word or character boundaries โ often breaks semantic units, leading to retrievals that lack context or contain partial information. Sophisticated chunking strategies preserve meaning, maintain overlap for context, and adapt to document structure (HTML, PDF, code, markdown, etc.).
This guide covers the full spectrum of chunking techniques: from simple fixed-size chunks with overlap, to recursive splitting that respects document hierarchy, to cutting-edge semantic chunking that groups semantically similar sentences, to agentic chunking that uses LLMs to intelligently parse documents. You'll learn when to use each strategy, how to implement them, and how to optimize chunk sizes for your specific use case.
What You'll Learn
Core Concepts
Understand the relationship between chunk size, overlap, and retrieval quality. Learn why arbitrary splits fail and what makes a good chunk.
5 Chunking Strategies
Master fixed-size, recursive, semantic, document-aware, and agentic chunking with complete implementations in Python.
Optimization & Tuning
Learn how to measure chunking quality, run A/B experiments, and find the optimal chunk size and strategy for your documents.
Production-Ready Code
Build complete RAG pipelines using LangChain, LlamaIndex, and custom implementations with best practices for scaling.
Prerequisites
Familiarity with Python, embeddings/vector databases, basic NLP concepts, and RAG system architecture. Experience with LangChain or LlamaIndex is helpful but not required.
Why Chunking Matters
Chunking is the first critical decision point in a RAG pipeline. Everything downstream โ embedding quality, retrieval relevance, answer generation โ depends on chunking decisions. Yet it's often treated as a trivial detail rather than a core engineering challenge.
The Impact: A Real Example
Consider indexing a 100-page technical manual with 50,000 tokens. Three different chunking approaches:
Recursive (document-aware)
Semantic (optimal params)
Retrieval accuracy on domain-specific Q&A benchmark
Why This Matters
Context Preservation
Semantic coherence must be preserved. A chunk breaking mid-sentence loses context and confuses embeddings. Good chunking keeps related information together.
Retrieval Precision
A 512-token chunk retrieved at position 200-712 may contain irrelevant content before position 300. Semantic boundaries ensure retrieved context is relevant.
Embedding Quality
Embeddings encode semantic meaning best when chunks are semantically coherent. Arbitrary splits degrade embedding signal and retrieval accuracy.
Cost & Latency
Smaller chunks = more total chunks = more embeddings to store/query. Larger chunks = fewer chunks but lower retrieval precision. Finding the sweet spot saves money.
The Chunking Decision Tree
Your choice of chunking strategy depends on your document type and use case:
Document Type?
Decide based on structure
Plain text โ Code โ HTML/Markdown โ PDF
Semantics Important?
Is topic coherence critical?
No โ Fixed-Size | Yes โ Semantic/Recursive
Latency/Cost?
Can you afford embeddings?
Speed Critical โ Recursive | High Budget โ Semantic
Final Strategy
Implement & optimize
Key Insight: There is no universal “best” chunking strategy. The optimal approach depends on your documents, your queries, your LLM, and your retrieval model. The solution is: measure, experiment, iterate.
Core Concepts and Terminology
Before diving into specific strategies, understand the key concepts that define chunking behavior.
1. Chunk Size
The number of tokens (or characters) in each chunk. Typical range: 128-2048 tokens.
Small Chunks (128-256 tokens)
Pros: High retrieval precision, less irrelevant context. Cons: May lack context, requires more embeddings, higher query latency.
Medium Chunks (512 tokens)
Pros: Balance of context and precision, industry standard. Cons: May miss long-range dependencies.
Large Chunks (1024+ tokens)
Pros: Rich context, fewer embeddings. Cons: Lower precision, may include irrelevant content, slower retrieval.
2. Chunk Overlap
The number of tokens repeated at the boundary between consecutive chunks. Preserves context and maintains continuity.
Why Overlap Matters
Overlap ensures that important information split across chunk boundaries doesn't get lost. Without overlap, a sentence split between chunk N and chunk N+1 might only appear in one chunk's embedding. With overlap, it appears in both, improving retrieval probability.
3. Semantic Coherence
A good chunk contains related, cohesive information. A bad chunk arbitrarily splits meaning.
Bad Split Example:
Chunk 1: "The transformer architecture introduced in 2017 has become the foundation of modern... [cuts at 512 tokens]"
Chunk 2: "...NLP systems. Self-attention enables parallel processing. Unlike RNNs which process sequentially, transformers can process entire..." [continues]
Good Split Example:
Chunk 1: "The transformer architecture, introduced in 2017 by Vaswani et al., replaced RNNs as the dominant sequence model. Self-attention enables parallel processing, making transformers dramatically faster to train."
Chunk 2: "Unlike RNNs which process tokens sequentially, transformers process entire sequences simultaneously. This parallelization, combined with layer normalization and residual connections, enables training on massive datasets..."
4. Metadata Preservation
Preserving source document metadata (title, section, URL, author) enables:
- Better answer attribution ("According to Section 4.2...")
- Filtering retrieved chunks by document source
- Tracing retrieval decisions back to source
- Evaluating which documents are most helpful
5. The Retrieval Quality Spectrum
Irrelevant (0-20%)
Retrieved content doesn't answer the query
Partially Relevant (20-60%)
Contains some useful info but mixed with irrelevant content
Relevant (60-85%)
Mostly relevant with some tangential content
Highly Relevant (85-100%)
Directly answers the query with minimal noise
The Chunking Challenge: You're optimizing for retrieval quality โ finding chunks that are relevant to a query โ but you won't know query intent until runtime. This is why choosing the right chunking strategy upfront is so important.
Fixed-Size Chunking with Overlap
The simplest chunking strategy: split documents into chunks of fixed size (e.g., 512 tokens) with optional overlap (e.g., 20% overlap = 102 tokens).
How It Works
Iterate through the document, creating chunks at regular intervals. When using overlap, the end of one chunk repeats at the start of the next.
def fixed_size_chunking(text, chunk_size=512, overlap=0.2):
"""
Split text into fixed-size chunks with overlap.
Args:
text: Input document text
chunk_size: Size of each chunk in tokens (approximate)
overlap: Overlap ratio (0.0 to 1.0)
Returns:
List of chunk strings
"""
words = text.split()
overlap_size = int(chunk_size * overlap)
step_size = chunk_size - overlap_size
chunks = []
for i in range(0, len(words), step_size):
chunk = words[i:i + chunk_size]
chunks.append(" ".join(chunk))
if i + chunk_size >= len(words):
break
return chunks
document = """
The transformer architecture revolutionized NLP.
Self-attention mechanisms enable parallel processing...
[... more text ...]
"""
chunks = fixed_size_chunking(document, chunk_size=128, overlap=0.2)
for i, chunk in enumerate(chunks):
print(f"Chunk {i}: {len(chunk.split())} words")
print(f" {chunk[:80]}...\n")
Advantages
- Simplicity: Easy to implement and understand
- Speed: No computation overhead, runs instantly
- Predictability: Chunk count is predictable (useful for cost estimation)
- Flexibility: Works with any document type
Disadvantages
- Semantic Ignorance: Doesn't respect paragraph/section boundaries
- Context Loss: May split sentences or ideas mid-way
- One-Size-Fits-All: Same strategy regardless of document structure
- Lower Retrieval Quality: Typically 60-70% retrieval accuracy
When to Use
Fixed-size chunking is appropriate when:
- You're prototyping quickly and don't need optimal quality
- Documents are highly homogeneous plain text
- You need to be language-agnostic (no tokenizer dependencies)
- You're optimizing purely for speed over quality
Exercise: Fixed-Size with Different Overlap
Implement fixed-size chunking with 0%, 20%, and 50% overlap on a sample document. Compare the number of chunks and overlap between consecutive chunks.
document = "Your text here..."
for overlap_pct in [0, 0.2, 0.5]:
chunks = fixed_size_chunking(document, 256, overlap_pct)
print(f"Overlap {overlap_pct*100}%: {len(chunks)} chunks")
Token Counting
The example uses word count as an approximation of token count. In production, use a proper tokenizer (tiktoken for GPT, sentencepiece for others) to match your embedding model's tokenization.
Recursive Chunking
Recursive chunking respects document structure by attempting to split on progressively more granular separators (paragraph โ sentence โ word โ character) until chunks fit the target size.
How It Works
The algorithm maintains a list of separator patterns (hierarchical). It tries the first separator; if chunks are still too large, it tries the next separator within those chunks, recursively, until all chunks fit the target size.
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(
chunk_size=512,
chunk_overlap=102,
separators=[
"\n\n",
"\n",
". ",
" ",
""
]
)
document = """
Chapter 1: Introduction
Paragraph 1 about transformers...
Paragraph 2 about attention...
"""
chunks = splitter.split_text(document)
for i, chunk in enumerate(chunks):
print(f"Chunk {i} ({len(chunk)} chars): {chunk[:60]}...")
Why Recursive Splitting?
Semantic Preservation
By preferring paragraph splits over sentence splits, recursive splitting keeps related paragraphs together when possible.
Hierarchy Awareness
The separator list encodes document structure: chapters > paragraphs > sentences > words. Respects this hierarchy when possible.
Flexibility
Automatically adapts. If paragraphs are too large, it splits sentences. If sentences are still too large, it splits words.
Separator Strategies for Different Document Types
markdown_seps = ["\n\n", "\n", "# ", "## ", "### ", ". ", " ", ""]
code_seps = ["\n\n", "\nclass ", "\ndef ", "\n\n", "\n", " ", ""]
html_seps = ["\n\n", "\n", "
", "", "", ". ", " ", ""]
# Plain Text (Default)
text_seps = ["\n\n", "\n", ". ", " ", ""]
Advantages
- Structure-Aware: Respects document hierarchy
- Better Semantics: Keeps paragraphs together when possible
- Flexible: Handles many document types with custom separators
- Good Retrieval Quality: 75-85% accuracy typical
- Fast: Still very fast, no embeddings needed
Disadvantages
- Separator Dependent: Quality depends on separator choice
- Document Type Specific: May need different separator lists for different content
- Still Not Semantic: Doesn't understand meaning, only structure
- Edge Cases: May struggle with irregular document formatting
Exercise: Recursive Chunking
Use RecursiveCharacterTextSplitter with a markdown document. Try different separator lists and observe how chunk boundaries change.
from langchain.text_splitter import RecursiveCharacterTextSplitter
markdown_text = '''
# Chapter 1
## Section 1.1
Content here...
'''
splitter = RecursiveCharacterTextSplitter(chunk_size=200)
chunks = splitter.split_text(markdown_text)
Semantic Chunking with Embeddings
Semantic chunking uses embeddings to measure semantic similarity between sentences. Chunks are boundaries where semantic similarity drops below a threshold, preserving semantic coherence.
How It Works
- Split document into sentences
- Generate embeddings for each sentence
- Compute cosine similarity between consecutive sentences
- Identify breakpoints where similarity drops below threshold
- Group sentences into chunks based on these breakpoints
import numpy as np
from sklearn.metrics.pairwise import cosine_similarity
import re
def semantic_chunking(text, embedding_func, threshold=0.5, max_size=512):
"""
Chunk text using semantic similarity.
Args:
text: Input document
embedding_func: Function that returns embeddings (e.g., from OpenAI)
threshold: Similarity threshold for chunk boundaries (0-1)
max_size: Maximum tokens per chunk
Returns:
List of chunks
"""
sentences = re.split(r'(?<=[.!?])\s+', text)
sentences = [s.strip() for s in sentences if s.strip()]
if len(sentences) == 0:
return []
embeddings = []
for sentence in sentences:
emb = embedding_func(sentence)
embeddings.append(emb)
similarities = []
for i in range(len(embeddings) - 1):
sim = cosine_similarity(
[embeddings[i]],
[embeddings[i+1]]
)[0][0]
similarities.append(sim)
breakpoints = [0]
for i, sim in enumerate(similarities):
if sim < threshold:
breakpoints.append(i + 1)
breakpoints.append(len(sentences))
chunks = []
for start, end in zip(breakpoints[:-1], breakpoints[1:]):
chunk = " ".join(sentences[start:end])
if len(chunk) > max_size:
chunks.extend(semantic_chunking(
chunk, embedding_func, threshold, max_size
))
else:
chunks.append(chunk)
return chunks
Example with OpenAI Embeddings
from openai import OpenAI
client = OpenAI()
def get_embedding(text):
response = client.embeddings.create(
input=text,
model="text-embedding-3-small"
)
return response.data[0].embedding
document = """
The transformer architecture has revolutionized AI.
It introduced self-attention mechanisms that enable
parallel processing. Unlike RNNs, transformers process
entire sequences at once. This architectural choice,
combined with scale, unlocked the modern AI era.
"""
chunks = semantic_chunking(document, get_embedding, threshold=0.6)
for i, chunk in enumerate(chunks):
print(f"Chunk {i}: {chunk[:70]}...")
Advantages
- Semantic Coherence: Each chunk is semantically cohesive
- Optimal Boundaries: Chunks break where topics naturally change
- High Retrieval Quality: 85-95% accuracy typical
- Universal: Works on any document type and language
- Principled: Grounded in meaningful similarity metrics
Disadvantages
- Computational Cost: Requires embeddings for every sentence (expensive)
- Latency: Much slower than fixed-size or recursive (seconds vs milliseconds)
- Threshold Tuning: Requires careful threshold selection (0.5 vs 0.7 yields very different results)
- Language Dependent: Quality depends on embedding model quality
Threshold Tuning
The similarity threshold dramatically affects chunk boundaries:
Threshold = 0.3
Many small chunks, low latency
Threshold = 0.5
Balanced chunks, balanced latency
Threshold = 0.8
Few large chunks, high latency
Exercise: Semantic Chunking
Implement semantic chunking with different thresholds (0.3, 0.5, 0.8). Compare the number of chunks and their sizes.
import numpy as np
def mock_embedding(text):
np.random.seed(hash(text) % 2**32)
return np.random.randn(384)
Document-Aware Chunking
Different document types (HTML, PDF, code, markdown) have distinct structure. Document-aware chunking preserves this structure, respecting hierarchy, formatting, and semantic units specific to each type.
HTML Document Chunking
from bs4 import BeautifulSoup
from langchain.document_loaders import UnstructuredHTMLLoader
html_content = """
<article>
<h1>Machine Learning Fundamentals</h1>
<section id="intro-2">
<h2>Introduction</h2>
<p>Machine learning is...</p>
<p>There are three main types...</p>
</section>
<section id="supervised">
<h2>Supervised Learning</h2>
<p>In supervised learning...</p>
</section>
</article>
"""
def chunk_html(html_text, chunk_size=512):
soup = BeautifulSoup(html_text, 'html.parser')
chunks = []
for section in soup.find_all(['section', 'article']):
heading = section.find(['h1', 'h2', 'h3'])
heading_text = heading.text if heading else ""
paragraphs = section.find_all('p')
content = heading_text + "\n" + "\n".join(p.text for p in paragraphs)
if len(content) > chunk_size:
words = content.split()
for i in range(0, len(words), chunk_size//10):
chunks.append(" ".join(words[i:i+chunk_size//10]))
else:
chunks.append(content)
return chunks
chunks = chunk_html(html_content)
for i, chunk in enumerate(chunks):
print(f"Chunk {i}: {chunk[:80]}...")
Code Document Chunking
import re
def chunk_code(code_text, chunk_size=512):
"""
Chunk code while preserving class and function boundaries.
"""
chunks = []
current_chunk = ""
blocks = re.split(r'^(class |def )', code_text, flags=re.MULTILINE)
for block in blocks:
if len(current_chunk) + len(block) <= chunk_size:
current_chunk += block
else:
if current_chunk:
chunks.append(current_chunk)
current_chunk = block
if current_chunk:
chunks.append(current_chunk)
return chunks
code = """
class DataProcessor:
def __init__(self):
self.data = []
def process(self, item):
return item.upper()
def helper_function():
return "result"
"""
chunks = chunk_code(code, chunk_size=100)
for i, chunk in enumerate(chunks):
print(f"Code Chunk {i}:\n{chunk}\n---")
PDF Document Chunking
from PyPDF2 import PdfReader
from langchain.document_loaders import PyPDFLoader
def chunk_pdf(pdf_path, chunk_size=512):
"""
Extract text from PDF preserving page and section structure.
"""
loader = PyPDFLoader(pdf_path)
documents = loader.load()
chunks = []
for doc in documents:
text = doc.page_content
metadata = doc.metadata
words = text.split()
for i in range(0, len(words), chunk_size):
chunk_text = " ".join(words[i:i+chunk_size])
chunks.append({
"text": chunk_text,
"page": metadata.get("page", 0),
"source": metadata.get("source", "")
})
return chunks
Metadata Enrichment
Document-aware chunking enables metadata enrichment, improving attribution and filtering:
chunk_with_metadata = {
"text": "The transformer architecture...",
"chunk_id": "doc_001_chunk_005",
"document": "learn-transformers.html",
"section": "Core Concepts",
"subsection": "Self-Attention",
"page": 3,
"start_char": 1250,
"end_char": 1762,
"language": "en",
"document_type": "html"
}
Advantages
- Structure Preservation: Maintains document hierarchy and formatting
- Better Attribution: Rich metadata enables precise source citation
- Specialized Handling: Code, HTML, PDF each get optimal treatment
- Filtering Capability: Can filter by section, page, document type, etc.
Disadvantages
- Implementation Complexity: Different logic for each document type
- Dependency Overhead: Requires specialized libraries (BeautifulSoup, PyPDF2, etc.)
- Error Prone: Parsing is brittle; formatting issues break chunking
Exercise: Multi-Format Chunking
Implement a chunking pipeline that handles markdown, HTML, and plain text, preserving structure for each format.
def universal_chunk(text, format_type, chunk_size=512):
if format_type == 'html':
return chunk_html(text, chunk_size)
elif format_type == 'markdown':
return chunk_markdown(text, chunk_size)
else:
return fixed_size_chunking(text, chunk_size)
Agentic Chunking
Agentic chunking uses an LLM as an agent to intelligently parse documents, deciding how to chunk based on semantic meaning rather than mechanical rules.
How It Works
The LLM reads a document and decides:
- Which content belongs together
- Where natural boundaries exist
- Which details are important vs contextual
- How to create chunks with necessary context
from anthropic import Anthropic
client = Anthropic()
def agentic_chunking(document_text, model="claude-3-5-sonnet-20241022"):
"""
Use Claude to intelligently chunk a document.
"""
response = client.messages.create(
model=model,
max_tokens=4096,
messages=[
{
"role": "user",
"content": f"""Analyze this document and identify natural semantic chunks.
For each chunk:
1. Define clear boundaries
2. Identify the main topic
3. List key concepts
4. Note any required context
Document:
{document_text}
Provide your response as structured JSON with chunks array."""
}
]
)
response_text = response.content[0].text
import json
import re
json_match = re.search(r'\{.*\}', response_text, re.DOTALL)
if json_match:
chunks_data = json.loads(json_match.group())
return chunks_data["chunks"]
return []
document = """
Chapter 1: Neural Networks
A neural network is a computational model...
Neurons process information...
Chapter 2: Training
To train a neural network, we use gradient descent...
Backpropagation computes gradients...
"""
chunks = agentic_chunking(document)
for i, chunk in enumerate(chunks):
print(f"Chunk {i}: {chunk.get('topic', 'Unknown')}")
print(f" Content: {chunk.get('text', '')[:100]}...")
Advanced: Multi-Pass Agentic Chunking
def multi_pass_agentic_chunking(document, max_passes=3):
"""
Iteratively refine chunking through multiple LLM passes.
Pass 1: Identify high-level structure
Pass 2: Refine boundaries based on semantic analysis
Pass 3: Verify chunk quality and optimize
"""
chunks = []
for pass_num in range(max_passes):
if pass_num == 0:
prompt = f"""Identify the main topics and high-level structure.
Document: {document}"""
elif pass_num == 1:
prompt = f"""Refine chunk boundaries to be semantically coherent.
Previous analysis: {chunks}"""
else:
prompt = f"""Verify chunk quality. Are chunks:
1. Semantically coherent?
2. Self-contained?
3. Complete (not mid-sentence)?
Current chunks: {chunks}"""
response = client.messages.create(
model="claude-3-5-sonnet-20241022",
max_tokens=2048,
messages=[{"role": "user", "content": prompt}]
)
return chunks
Advantages
- Intelligent Parsing: Uses semantic understanding, not rules
- Highest Quality: Often 95%+ retrieval accuracy
- Context-Aware: Creates chunks with all needed context
- Flexible: Can follow any custom chunking criteria you specify
- Explainable: LLM can explain why chunks are structured that way
Disadvantages
- Cost: LLM API calls are expensive (significant budget impact at scale)
- Latency: Seconds to minutes for large documents
- Non-Deterministic: Results vary across runs (temperature dependency)
- Token Limits: Can't handle extremely large documents in one pass
- Context Window: Limited by model context window size
Cost Analysis
Semantic (embeddings)
$0.02-0.05 per 100K tokens
Agentic (Claude)
$0.30-1.00 per 100K tokens
Approximate cost for processing 100K tokens
When to use Agentic Chunking: When document quality is critical (legal documents, medical records), or when documents are complex and standard strategies fail. Not suitable for high-volume processing with tight budgets.
Chunking Strategy Comparison
Here's a comprehensive comparison of all five strategies across key dimensions:
| Dimension |
Fixed-Size |
Recursive |
Semantic |
Doc-Aware |
Agentic |
| Retrieval Accuracy |
62% |
78% |
91% |
85% |
96% |
| Speed |
Instant (ms) |
Fast (ms) |
Moderate (secs) |
Moderate (secs) |
Slow (mins) |
| Cost |
Free |
Free |
$0.02/100K |
$0.02/100K |
$0.50/100K |
| Implementation Complexity |
Very Easy |
Easy |
Moderate |
Moderate |
Complex |
| Deterministic |
Yes |
Yes |
Yes |
Yes |
No (LLM variance) |
| Best For |
Prototyping, simple documents |
Structured text, markdown |
Mixed documents, high quality |
Specific formats (HTML, PDF) |
Critical documents, best quality |
Decision Matrix
Use this matrix to choose the right strategy for your use case:
Choose Fixed-Size if:
- Building a rapid prototype
- Need zero overhead
- Documents are homogeneous
Choose Recursive if:
- Documents are structured (markdown, code)
- Need decent quality without cost
- Want simple implementation
Choose Semantic if:
- High retrieval quality needed
- Have embedding budget
- Mixed document types
Choose Agentic if:
- Critical documents (legal, medical)
- Quality > cost/latency
- Complex domain knowledge needed
Chunk Size Optimization Experiment
The optimal chunk size is not universal โ it depends on your documents, queries, retrieval model, and LLM. The best approach is empirical: test different sizes and measure actual retrieval quality.
Experiment Setup
Here's a complete framework for optimizing chunk size:
import json
import time
from sklearn.metrics import ndcg_score
from tqdm import tqdm
def run_chunk_optimization(documents, queries, chunk_sizes, embedding_func):
"""
Test different chunk sizes and measure retrieval quality.
Returns:
Dictionary of results by chunk size
"""
results = {}
for chunk_size in chunk_sizes:
print(f"\nTesting chunk_size={chunk_size}")
start_time = time.time()
all_chunks = []
for doc in documents:
chunks = semantic_chunking(doc, embedding_func, max_size=chunk_size)
all_chunks.extend(chunks)
chunking_time = time.time() - start_time
embeddings = []
for chunk in all_chunks:
emb = embedding_func(chunk)
embeddings.append(emb)
ndcg_scores = []
for query in tqdm(queries, desc=f"Evaluating {len(queries)} queries"):
query_emb = embedding_func(query["text"])
similarities = [cosine_similarity([query_emb], [emb])[0][0]
for emb in embeddings]
ranked_indices = sorted(range(len(similarities)),
key=lambda i: similarities[i],
reverse=True)[:5]
relevances = [1 if idx in query["relevant_chunk_ids"] else 0
for idx in ranked_indices]
ndcg = ndcg_score([query["relevances"]], [relevances])
ndcg_scores.append(ndcg)
results[chunk_size] = {
"avg_ndcg": sum(ndcg_scores) / len(ndcg_scores),
"num_chunks": len(all_chunks),
"chunking_time": chunking_time,
"chunks_per_query": len(all_chunks) / len(queries)
}
return results
chunk_sizes = [128, 256, 512, 1024, 2048]
results = run_chunk_optimization(documents, test_queries, chunk_sizes, get_embedding)
print("\n=== Optimization Results ===")
for size in chunk_sizes:
r = results[size]
print(f"Size {size:4d}: NDCG={r['avg_ndcg']:.3f}, "
f"Chunks={r['num_chunks']:4d}, "
f"Time={r['chunking_time']:.2f}s")
Key Metrics to Track
NDCG@k Score
Normalized Discounted Cumulative Gain measures ranking quality. Score from 0-1, higher is better. Standard in IR evaluation.
MRR (Mean Reciprocal Rank)
Average position of the first relevant chunk. Measures how quickly you find relevant results.
Recall@k
Percentage of relevant chunks in top-k results. Critical for RAG (need relevant info to generate good answers).
Latency
Time to chunk documents and query the index. Includes embedding time and database queries.
Typical Results from Optimization
Typical NDCG@5 scores by chunk size (results vary by domain)
Optimization Insights
Sweet Spot
In most cases, 256-512 tokens yields the best NDCG scores. Smaller chunks miss context; larger chunks include noise. Your optimization should confirm this for your specific data.
Domain Sensitivity
Technical documents benefit from larger chunks (512+) to maintain context. News articles and short documents benefit from smaller chunks (128-256). Always test on representative data from your domain.
Overlap Impact
Overlap improves recall at the cost of more total chunks. 20-25% overlap typically provides good balance. Too much overlap (>50%) creates excessive redundancy without quality gains.
Retrieval Quality by Chunk Size
This section visualizes how chunk size affects the three critical dimensions of retrieval quality:
1. Precision vs Recall Tradeoff
2. Latency by Chunk Size and Strategy
3. Cost Analysis: Chunking + Embedding + Storage
Scenario: Process 1 million documents (100 tokens each) and serve 1000 queries/day
Fixed-Size
Chunking: Free
Embedding: N/A
Storage: 1M chunks
Total: $0
Semantic
Chunking: Free
Embedding: 100M tokens โ $20
Storage: 1M vectors
Total: $20 (+$5/mo)
Agentic
Chunking: 100M tokens โ $300
Embedding: 1M chunks โ $20
Storage: 1M vectors
Total: $320
Advanced Chunking Techniques
1. Sliding Window Chunking
Maximize context by overlapping chunks more intelligently. Instead of uniform overlap, use sliding windows with variable step sizes.
def sliding_window_chunking(text, window_size=512, step_size=256):
"""
Create overlapping chunks using sliding window approach.
Args:
text: Input document
window_size: Size of each window
step_size: How far to move between windows
Returns:
List of chunks with indices
"""
words = text.split()
chunks_with_info = []
for start in range(0, len(words) - window_size, step_size):
chunk = " ".join(words[start:start + window_size])
chunks_with_info.append({
"text": chunk,
"start_idx": start,
"end_idx": start + window_size,
"overlap_with_next": window_size - step_size
})
return chunks_with_info
2. Hierarchical Chunking
Create a hierarchy of chunks: summary chunks, section chunks, paragraph chunks. At query time, start with summary chunks to narrow scope, then retrieve detailed chunks.
def hierarchical_chunking(document):
"""
Create multi-level chunk hierarchy.
"""
hierarchy = {
"document_summary": summarize(document),
"sections": [],
}
for section in document.sections:
section_data = {
"title": section.title,
"summary": summarize(section.text),
"paragraphs": []
}
for paragraph in section.paragraphs:
section_data["paragraphs"].append({
"text": paragraph,
"summary": summarize(paragraph)
})
hierarchy["sections"].append(section_data)
return hierarchy
3. Colbert Chunking (Contextual Dense Embedding)
ColBERT uses late interaction for dense retrieval. Instead of embedding entire chunks, embed individual tokens and compute interaction at query time. Enables finer-grained chunking without quality loss.
4. Graph-Based Chunking
Represent documents as knowledge graphs and chunk based on semantic graph structure rather than linear text order.
Node Parsers and LlamaIndex Integration
LlamaIndex provides production-grade node parsers that abstract chunking complexity. Nodes are the LlamaIndex term for chunks โ they include the text, metadata, and embedding vectors.
Simple Node Parser
from llama_index.core.node_parser import SimpleNodeParser
parser = SimpleNodeParser.from_defaults(
chunk_size=512,
chunk_overlap=102
)
from llama_index.core import Document
docs = [
Document(text="Your document text here..."),
Document(text="Another document...")
]
nodes = parser.get_nodes_from_documents(docs)
for node in nodes:
print(f"Node: {node.node_id}")
print(f"Text: {node.text[:100]}...")
print(f"Metadata: {node.metadata}")
print()
Semantic Chunking with LlamaIndex
from llama_index.core.node_parser import SemanticSplitterNodeParser
from llama_index.embeddings.openai import OpenAIEmbedding
embed_model = OpenAIEmbedding(model="text-embedding-3-small")
parser = SemanticSplitterNodeParser(
buffer_size=1,
breakpoint_percentile_threshold=95,
embed_model=embed_model,
)
nodes = parser.get_nodes_from_documents(docs)
Hierarchical Node Parser
from llama_index.core.node_parser import HierarchicalNodeParser
parser = HierarchicalNodeParser.from_defaults(
chunk_size=512,
chunk_overlap=102,
)
nodes = parser.get_nodes_from_documents(docs)
for node in nodes:
if hasattr(node, 'is_parent') and node.is_parent:
print(f"Summary Node: {node.text[:80]}...")
else:
print(f"Leaf Node: {node.text[:80]}...")
Integration with RAG Pipeline
from llama_index.core import VectorStoreIndex
parser = SemanticSplitterNodeParser(
buffer_size=1,
breakpoint_percentile_threshold=95,
embed_model=embed_model,
)
nodes = parser.get_nodes_from_documents(docs)
index = VectorStoreIndex(nodes)
query_engine = index.as_query_engine()
response = query_engine.query("How do transformers work?")
print(response)
Node Properties and Metadata
{
"node_id": "abc123",
"text": "The chunk content...",
"metadata": {
"page_label": "1",
"file_name": "document.pdf",
"doc_id": "doc_001",
"section": "Introduction",
"start_char_idx": 0,
"end_char_idx": 512
},
"relationships": {
"parent": "parent_node_id",
"next": "next_node_id",
"prev": "prev_node_id"
},
"embedding": [0.1, 0.2, ..., 0.9]
}
Custom Node Parser
from llama_index.core.node_parser import NodeParser
from llama_index.core.schema import BaseNode
class CustomNodeParser(NodeParser):
def _parse_nodes(self, docs, **kwargs):
"""Implement your custom chunking logic."""
nodes = []
for doc in docs:
chunks = your_chunking_function(doc.text)
for i, chunk_text in enumerate(chunks):
node = BaseNode(
text=chunk_text,
metadata={
"chunk_idx": i,
"doc_id": doc.doc_id,
**doc.metadata
}
)
nodes.append(node)
return nodes
parser = CustomNodeParser()
nodes = parser.get_nodes_from_documents(docs)
Best Practices for Production Chunking
1. Always Test on Representative Data
Create a test set of documents and queries that reflect your actual use case. Chunking decisions must be validated on real data, not synthetic examples.
Best Practice
Use your production queries to evaluate chunking quality. Run A/B tests with different chunk sizes and strategies. Track NDCG, MRR, and Recall@5. Choose the strategy with best overall quality/cost tradeoff.
2. Preserve Document Structure and Metadata
chunk = {
"text": "The chunk content...",
"source_doc": "technical_guide.pdf",
"page": 5,
"section": "3.2 Architecture",
"timestamp": "2024-02-27",
"language": "en",
"quality_score": 0.92
}
3. Implement Versioning for Chunks
When you update chunking strategy, version your chunks. This allows rolling back if a new strategy performs worse.
chunk = {
"chunk_id": "v2_doc_001_chunk_005",
"version": 2,
"strategy": "semantic-0.6-threshold",
"text": "...",
"created_at": "2024-02-27T10:30:00Z"
}
4. Monitor Chunk Quality Metrics
Track these metrics continuously in production:
Retrieval Quality
NDCG@5, MRR, Recall@5. Should remain consistent over time.
Chunk Size Distribution
Mean, median, p95 chunk sizes. Should match your target.
Answer Quality
BLEU, ROUGE scores of final LLM answers vs gold answers.
Latency
End-to-end latency from query to retrieved chunks. Should be <100ms.
5. Handle Edge Cases
- Implement hierarchical chunking
- Or split into subdocuments first
- Don't chunk, use entire document as one chunk
- Or add padding/context from related documents
# Tables and structured data
- Parse tables separately, include col headers in chunks
- Don't split rows arbitrarily
- Respect bracket/indentation structure
- Include full function/class definitions, not partial
- Use language-specific tokenizers
- May need different chunk sizes per language
6. Optimize for Your Retrieval Model
The best chunk size depends on your retrieval model:
- BM25 (lexical search): 256-512 tokens. Respects word frequency patterns.
- Dense embeddings (semantic): 256-512 tokens. Optimal for sentence transformers.
- ColBERT (late interaction): 128-256 tokens. Can work with shorter chunks.
- Hybrid (BM25 + semantic): 512 tokens. Larger chunks benefit both methods.
7. Implement Chunk Quality Filtering
def should_keep_chunk(chunk_text, min_words=5, max_redundancy=0.3):
"""Filter out chunks that don't meet quality thresholds."""
if len(chunk_text.split()) < min_words:
return False
if len(chunk_text.strip()) < len(chunk_text) * 0.5:
return False
alphanumeric = sum(1 for c in chunk_text if c.isalnum())
if alphanumeric < len(chunk_text) * 0.3:
return False
sentences = chunk_text.split(". ")
if len(sentences) > 0:
unique_ratio = len(set(sentences)) / len(sentences)
if unique_ratio < (1 - max_redundancy):
return False
return True
quality_chunks = [c for c in chunks if should_keep_chunk(c)]
Common Mistakes and How to Avoid Them
Mistake 1: Using Fixed-Size Chunking for Everything
Problem: Fixed-size ignores document structure. A 512-token chunk may cut paragraphs and concepts mid-way.
Solution: Use recursive or semantic chunking for heterogeneous documents. Fixed-size only for homogeneous plain text.
Mistake 2: No Overlap Between Chunks
Problem: Information split between chunks may be missed. Queries aligned with chunk boundaries may lose relevant context.
Solution: Always use 20-25% overlap. The ~2x cost in total chunks is worth the retrieval quality improvement.
Mistake 3: Choosing Chunk Size Without Testing
Problem: Arbitrary choices (e.g., "use 1024 tokens") ignore your specific data and use case.
Solution: Run optimization experiments as shown in Section 10. Test 128, 256, 512, 1024 on representative data.
Mistake 4: Losing Source Attribution
Problem: If chunks don't preserve document/page metadata, you can't cite sources in answers.
Solution: Always include source_doc, page, section, and character offsets in chunk metadata. Pass this through to final answers.
Mistake 5: Not Handling Document-Specific Formats
Problem: Using generic chunking on HTML, PDF, or code loses structure. HTML tags get included in embeddings. PDF page breaks aren't preserved.
Solution: Use document-aware parsers. Use LangChain loaders and LlamaIndex node parsers that handle format-specific structure.
Mistake 6: Using Inconsistent Tokenization
Problem: Word count โ token count. If you chunk by word count but your embedding model uses byte-pair encoding (BPE), actual token count may be 1.3x higher.
Solution: Use the same tokenizer as your embedding model. Use tiktoken for OpenAI models, use HF transformers tokenizers for others.
Mistake 7: Semantic Chunking Without Threshold Tuning
Problem: The default threshold (0.5) may not be optimal for your documents. Using it without tuning wastes the semantic approach's potential.
Solution: Analyze similarity distributions in your documents. Test thresholds from 0.3 to 0.8. Pick the one with best NDCG on test queries.
Mistake 8: Forgetting to Re-Chunk When Documents Update
Problem: If a document is updated but chunks aren't regenerated, your index becomes stale and questions about updated content get wrong answers.
Solution: Implement document versioning and incremental chunking. Detect updated documents and re-chunk only those. Invalidate old embeddings.
Complete Code Examples and Recipes
Recipe 1: End-to-End RAG with Optimal Chunking
from llama_index.core import VectorStoreIndex, Document
from llama_index.core.node_parser import SemanticSplitterNodeParser
from llama_index.embeddings.openai import OpenAIEmbedding
from llama_index.llms.openai import OpenAI
import os
embed_model = OpenAIEmbedding(model="text-embedding-3-small")
llm = OpenAI(model="gpt-4-turbo", temperature=0.7)
documents = [
Document(text="Your document content...",
metadata={"source": "file1.pdf", "page": 1}),
Document(text="Another document...",
metadata={"source": "file2.pdf", "page": 1}),
]
parser = SemanticSplitterNodeParser(
buffer_size=1,
breakpoint_percentile_threshold=95,
embed_model=embed_model,
)
nodes = parser.get_nodes_from_documents(documents)
index = VectorStoreIndex(nodes)
query_engine = index.as_query_engine(
llm=llm,
similarity_top_k=5,
)
response = query_engine.query("What is semantic chunking?")
print(response)
Recipe 2: Chunking with Quality Filtering
def chunk_with_quality_filter(text, chunk_size=512, overlap=0.2, min_quality=0.7):
"""Chunk text and filter by quality."""
chunks = semantic_chunking(text, chunk_size=chunk_size, overlap=overlap)
quality_chunks = []
for chunk in chunks:
score = compute_quality_score(chunk)
if score >= min_quality:
quality_chunks.append({
"text": chunk,
"quality_score": score,
"length": len(chunk.split())
})
return quality_chunks
def compute_quality_score(chunk):
"""Estimate chunk quality (0-1)."""
score = 0
if chunk.endswith((".","!","?")):
score += 0.2
sentences = len(chunk.split("."))
if 2 <= sentences <= 5:
score += 0.3
unique_words = len(set(chunk.lower().split()))
total_words = len(chunk.split())
if unique_words / total_words > 0.6:
score += 0.5
return min(score, 1.0)
Recipe 3: Multi-Format Document Processing
from pathlib import Path
from PyPDF2 import PdfReader
from bs4 import BeautifulSoup
def chunk_any_document(file_path, chunk_size=512):
"""Intelligently chunk any document type."""
path = Path(file_path)
if path.suffix == ".pdf":
text = extract_text_from_pdf(file_path)
elif path.suffix == ".html":
text = extract_text_from_html(file_path)
elif path.suffix == ".md":
text = extract_text_from_markdown(file_path)
else:
text = path.read_text()
chunks = semantic_chunking(text, chunk_size=chunk_size)
return [
{"text": chunk, "source": str(file_path), "format": path.suffix}
for chunk in chunks
]
def extract_text_from_pdf(pdf_path):
reader = PdfReader(pdf_path)
text = ""
for page in reader.pages:
text += page.extract_text()
return text
def extract_text_from_html(html_path):
soup = BeautifulSoup(open(html_path).read(), "html.parser")
for tag in soup(["script", "style"]):
tag.decompose()
return soup.get_text()
def extract_text_from_markdown(md_path):
return open(md_path).read()
Hands-On Exercises
Exercise 1: Implement Fixed-Size Chunking
Goal: Build a fixed-size chunker with overlap from scratch.
Fixed-Size Implementation
Implement fixed_size_chunking() with parameters for chunk_size and overlap. Test with different overlap values (0%, 20%, 50%) and verify the overlapping content.
def fixed_size_chunking(text, chunk_size=512, overlap=0.2):
words = text.split()
overlap_size = int(chunk_size * overlap)
step_size = chunk_size - overlap_size
chunks = []
return chunks
doc = "The transformer architecture...\n[100+ words]"
chunks = fixed_size_chunking(doc, 256, 0.2)
print(f"Created {len(chunks)} chunks")
if len(chunks) > 1:
pass
Exercise 2: Semantic Chunking with Threshold Tuning
Goal: Implement semantic chunking and find optimal threshold for a document.
Semantic Threshold Optimization
Implement semantic_chunking() with variable threshold. Test thresholds 0.3, 0.5, 0.7, 0.9 and plot: number of chunks vs. threshold.
from sklearn.metrics.pairwise import cosine_similarity
def semantic_chunking(text, embedding_func, threshold=0.5):
pass
thresholds = [0.3, 0.5, 0.7, 0.9]
chunk_counts = []
for threshold in thresholds:
chunks = semantic_chunking(document, get_embedding, threshold)
chunk_counts.append(len(chunks))
print(f"Threshold {threshold}: {len(chunks)} chunks")
import matplotlib.pyplot as plt
plt.plot(thresholds, chunk_counts)
plt.xlabel("Similarity Threshold")
plt.ylabel("Number of Chunks")
plt.show()
Exercise 3: Compare Chunking Strategies
Goal: Implement all 5 strategies and compare on a real document.
Strategy Comparison
Implement fixed_size_chunking, recursive_chunking, semantic_chunking, document_aware_chunking, and agentic_chunking. For each, measure: number of chunks, average chunk size, time to chunk.
import time
strategies = {
"fixed": (fixed_size_chunking, {}),
"recursive": (recursive_chunking, {}),
"semantic": (semantic_chunking, {"embedding_func": get_embedding}),
"doc_aware": (document_aware_chunking, {"doc_type": "html"}),
}
results = {}
for name, (func, kwargs) in strategies.items():
start = time.time()
chunks = func(document, **kwargs)
elapsed = time.time() - start
results[name] = {
"num_chunks": len(chunks),
"avg_size": sum(len(c.split()) for c in chunks) / len(chunks),
"time_sec": elapsed
}
print(f"{name:12}: {len(chunks):3} chunks, "
f"avg {results[name]['avg_size']:6.1f} words, "
f"time {elapsed:.3f}s")
Exercise 4: Chunk Quality Analysis
Goal: Develop metrics to measure chunk quality.
Quality Metrics
Compute for each chunk: coherence (do sentences relate to each other?), completeness (do sentences end properly?), length distribution. Identify low-quality chunks.
def analyze_chunk_quality(chunks):
metrics = {}
for i, chunk in enumerate(chunks):
is_complete = chunk.strip().endswith(('.', '!', '?'))
sentences = chunk.split('.')
words = len(chunk.split())
is_reasonable = 50 < words < 1000
metrics[i] = {
"complete": is_complete,
"coherence": 0.8,
"reasonable": is_reasonable,
"quality": 0.9
}
return metrics
quality = analyze_chunk_quality(chunks)
low_quality = [i for i, m in quality.items() if m["quality"] < 0.7]
print(f"Low quality chunks: {len(low_quality)} out of {len(chunks)}")
Exercise 5: Production RAG Pipeline
Goal: Build a complete RAG system with optimized chunking and measure end-to-end quality.
Production RAG Setup
Build complete pipeline: load PDFs โ chunk with your best strategy โ embed โ create vector index โ evaluate on test queries using NDCG@5 and MRR metrics.
from llama_index.core import VectorStoreIndex, Document
documents = load_pdf_directory("./documents/")
parser = SemanticSplitterNodeParser(...)
nodes = parser.get_nodes_from_documents(documents)
index = VectorStoreIndex(nodes)
test_queries = [
{"text": "What is X?", "relevant_doc_ids": [1, 3, 5]},
{"text": "How does Y work?", "relevant_doc_ids": [2, 4]},
]
ndcg_scores = []
for query in test_queries:
results = index.retrieve(query["text"], top_k=5)
ndcg_scores.append(ndcg)
print(f"Average NDCG@5: {sum(ndcg_scores)/len(ndcg_scores):.3f}")
Interview Questions on Chunking
Q1: Explain the tradeoff between chunk size and retrieval quality. ▼
Small chunks (128 tokens) have higher precision โ fewer irrelevant words per chunk. But they may lose context and require more embedding computations. Large chunks (1024+ tokens) preserve context but include noise. The sweet spot is usually 256-512 tokens, but this depends on your embedding model and query patterns. The only way to know is to test on your data.
Q2: Why is overlap important in chunking? ▼
Overlap ensures that information split across chunk boundaries appears in multiple chunks. Without overlap, a question aligned with a chunk boundary might miss relevant information. With 20% overlap, important facts are represented in both the boundary chunk and the previous chunk, improving retrieval probability. The cost is ~2x more total chunks, but retrieval quality improvement justifies it.
Q3: Compare fixed-size vs. recursive chunking. When would you use each? ▼
Fixed-size is simplest and fastest, good for prototyping. Recursive respects document structure (paragraphs > sentences > words), better for heterogeneous documents. Use fixed-size for homogeneous plain text or rapid prototyping. Use recursive for documentation, code, markdown โ anything with natural hierarchy. Recursive adds negligible latency overhead.
Q4: What are the advantages and costs of semantic chunking? ▼
Semantic chunking uses embeddings to find natural topic boundaries, yielding 85-95% retrieval accuracy vs 75-85% for recursive. Cost: must embed every sentence (expensive at scale, adds latency). Best for mixed documents where structure isn't reliable. Worth it if retrieval quality is critical. Not worth it for simple documents with clear structure.
Q5: How would you handle chunking for different document types (PDF, HTML, code)? ▼
Different document types have different structure: PDF has pages, HTML has DOM hierarchy, code has classes/functions. Use format-specific parsers (PyPDF2 for PDF, BeautifulSoup for HTML) to extract text while preserving structure. Then apply chunking that respects that structure. LangChain and LlamaIndex provide document loaders that handle this. Always preserve metadata (page, section, line number) in chunks.
Q6: What is agentic chunking and when would you use it? ▼
Agentic chunking uses an LLM to read documents and decide how to chunk them based on semantic meaning. Highest quality (95%+ retrieval accuracy) but expensive ($0.30-1.00 per 100K tokens) and slow (seconds to minutes). Use only when document quality is critical (legal, medical, financial) and cost/latency are acceptable. Not for high-volume processing with tight budgets.
Q7: You're building a RAG system and notice retrieval quality is 65%. How would you debug and improve it? ▼
First: measure chunking quality independently. Retrieve chunks for a test query and manually score relevance. If chunks are relevant but LLM answer is poor, problem is LLM, not chunking. If chunks are irrelevant: (1) test different chunk sizes via NDCG experiments, (2) switch to semantic chunking if using fixed-size, (3) ensure metadata is preserved so filtering works. Run A/B tests with different strategies and pick the winner.
Q8: Explain how you would optimize chunk size for a specific use case. ▼
Run experiments: test chunk sizes 128, 256, 512, 1024, 2048. For each: create chunks, embed all chunks, evaluate on 50-100 representative test queries. Measure NDCG@5, MRR, Recall@5. Also measure latency and cost (num chunks * embedding cost). Plot NDCG vs chunk size and pick the size with best NDCG, or best quality/cost tradeoff. Typical sweet spot: 256-512 tokens.
Frequently Asked Questions
What's the difference between chunking and tokenization? ▼
Tokenization breaks text into tokens (subword units), usually for LLMs. A token might be a word like 'transformer' or a subword like 're', 'form', 'ing'. Chunking creates semantically coherent units (sentences, paragraphs, sections) that may contain hundreds of tokens. Chunking happens before indexing; tokenization happens inside the embedding/LLM model.
Does chunk size affect embedding quality? ▼
Yes. Embedding models (sentence transformers, OpenAI embeddings) are optimized for certain input lengths. text-embedding-3-small works best with 256-512 tokens. If you send much shorter text, it may pad and be less discriminative. If you send much longer text (>2000 tokens), it may lose semantic coherence. Always test your embedding model's optimal input length.
How much should chunks overlap? ▼
20-25% overlap is standard. This means a 512-token chunk overlaps 102-128 tokens with the previous chunk. This ensures context continuity at boundaries. Too little overlap (<10%) loses context. Too much overlap (>50%) creates excessive redundancy without quality gains. The right amount depends on your data โ test and measure.
Should I chunk before or after extracting metadata? ▼
Extract and attach metadata first, then chunk. This ensures each chunk knows its source document, page, section, etc. If you chunk first, you lose track of structure. Process: parse document โ extract metadata (title, sections, page numbers) โ chunk while preserving metadata โ embed chunks โ index.
How do I handle documents that are 100K tokens long? ▼
Don't try to process in one pass. Split into smaller subdocuments first (e.g., by chapter or section). Chunk each subdocument independently. Preserve the hierarchy in metadata. Alternatively, use hierarchical chunking: create summary chunks at document level, then detailed chunks at paragraph level. At query time, use multi-stage retrieval: search summaries first, then detailed chunks.
Can I chunk documents in different languages differently? ▼
Yes, and you should. Languages have different average word lengths (Chinese characters vs English words), different punctuation conventions, different grammar. For multilingual documents: detect language per document, use language-specific tokenizers, potentially use different chunk sizes per language. LangChain and LlamaIndex support this with language-aware splitters.
How often should I re-chunk documents? ▼
When documents are updated, re-chunk them. When you change chunking strategy, re-chunk everything (versioning important). If using semantic chunking with embeddings, re-chunking also means re-embedding (cost!). Common approach: chunk documents once when ingested, version the chunks. When strategy changes, create v2, v3, etc. At query time, you can A/B test different versions.
What's the relationship between chunk size and context window for the LLM? ▼
If your LLM has a 4K token context window and you retrieve 5 chunks of 512 tokens, that's 2560 tokens of context plus prompt overhead. Chunks leave room for prompt instructions and the user question. Don't make chunks so large that retrieved chunks consume most of the context window. Leave at least 25% of context window for LLM to reason.
Should I include chunk overlap in the embedding, or just in the index? ▼
Include overlap in the index. When you embed chunks, embed the full chunk including the overlap. This is correct because overlapping sentences should be represented in both embeddings. The overlap isn't a separate embedding โ it's part of both chunks' embeddings.
How do I test if my chunking is good? ▼
Measure retrieval quality: run test queries, retrieve top-5 chunks, score relevance (1 = relevant, 0 = irrelevant), compute NDCG@5. Target >0.85 NDCG. Also track: MRR (how fast you find first relevant), Recall@5 (what fraction of relevant chunks appear), latency, cost. Collect human judgments on a small sample to calibrate automated metrics.
Summary and Next Steps
Key Takeaways
Chunking is Foundational
It's the first step in RAG. Small chunking mistakes compound through embedding, retrieval, and generation. Get it right.
5 Main Strategies
Fixed-size (simple), Recursive (structured), Semantic (high-quality), Document-Aware (format-specific), Agentic (best quality). Choose based on your use case.
Test Your Assumptions
The 'best' chunk size isn't universal. Test 128, 256, 512, 1024 on your data. Measure NDCG@5 and MRR. Use winners.
Preserve Everything
Keep metadata: source document, page, section, character offsets. This enables proper attribution and filtering.
Decision Framework: Which Strategy?
Fast Prototyping?
Use Fixed-Size. Zero overhead, instant results. Good enough for MVP.
Structured Documents?
Use Recursive. Respects hierarchy. Fast and effective.
High Quality Needed?
Use Semantic. Embedding cost worth retrieval gains.
Critical Documents?
Use Agentic. LLM understands complex semantics.
Next Steps
- Start with Recursive: Good default for most use cases. Structure-aware, fast, free.
- Benchmark Your Data: Run optimization experiments. Test chunk sizes 128-2048. Pick winner by NDCG.
- Implement Evaluation: Create test query set with human-judged relevance. Measure NDCG@5, MRR continuously.
- Consider Semantic: If recursive NDCG < 0.80, upgrade to semantic chunking with threshold tuning.
- Production Ready: Implement versioning, monitoring, re-chunking for updates. Track retrieval quality in production.
- Optimize Iteratively: As you gather production data, re-run optimization. Chunk size needs may change.
Further Resources
Final Insight: Chunking isn't a solved problem. It's domain-specific, data-specific, and model-specific. The teams building the best RAG systems invest heavily in chunking optimization because it's one of the highest-ROI improvements available. Even small NDCG gains (0.80 โ 0.85) significantly improve answer quality and user satisfaction.