Contents

Chunking Strategies for RAG Systems

Chunking is the foundational step in building effective Retrieval-Augmented Generation (RAG) systems. It's the process of intelligently splitting large documents into smaller, semantically coherent pieces that can be embedded, indexed, and retrieved to answer user queries. Done well, chunking dramatically improves retrieval accuracy and answer quality. Done poorly, it degrades both dramatically.

The naive approach โ€” splitting documents on arbitrary word or character boundaries โ€” often breaks semantic units, leading to retrievals that lack context or contain partial information. Sophisticated chunking strategies preserve meaning, maintain overlap for context, and adapt to document structure (HTML, PDF, code, markdown, etc.).

This guide covers the full spectrum of chunking techniques: from simple fixed-size chunks with overlap, to recursive splitting that respects document hierarchy, to cutting-edge semantic chunking that groups semantically similar sentences, to agentic chunking that uses LLMs to intelligently parse documents. You'll learn when to use each strategy, how to implement them, and how to optimize chunk sizes for your specific use case.

What You'll Learn

Core Concepts

Understand the relationship between chunk size, overlap, and retrieval quality. Learn why arbitrary splits fail and what makes a good chunk.

5 Chunking Strategies

Master fixed-size, recursive, semantic, document-aware, and agentic chunking with complete implementations in Python.

Optimization & Tuning

Learn how to measure chunking quality, run A/B experiments, and find the optimal chunk size and strategy for your documents.

Production-Ready Code

Build complete RAG pipelines using LangChain, LlamaIndex, and custom implementations with best practices for scaling.

Prerequisites

Familiarity with Python, embeddings/vector databases, basic NLP concepts, and RAG system architecture. Experience with LangChain or LlamaIndex is helpful but not required.

Why Chunking Matters

Chunking is the first critical decision point in a RAG pipeline. Everything downstream โ€” embedding quality, retrieval relevance, answer generation โ€” depends on chunking decisions. Yet it's often treated as a trivial detail rather than a core engineering challenge.

The Impact: A Real Example

Consider indexing a 100-page technical manual with 50,000 tokens. Three different chunking approaches:

Fixed-Size (512 tokens)
62% Retrieval Quality
Recursive (document-aware)
78% Retrieval Quality
Semantic (optimal params)
91% Retrieval Quality

Retrieval accuracy on domain-specific Q&A benchmark

Why This Matters

Context Preservation

Semantic coherence must be preserved. A chunk breaking mid-sentence loses context and confuses embeddings. Good chunking keeps related information together.

Retrieval Precision

A 512-token chunk retrieved at position 200-712 may contain irrelevant content before position 300. Semantic boundaries ensure retrieved context is relevant.

Embedding Quality

Embeddings encode semantic meaning best when chunks are semantically coherent. Arbitrary splits degrade embedding signal and retrieval accuracy.

Cost & Latency

Smaller chunks = more total chunks = more embeddings to store/query. Larger chunks = fewer chunks but lower retrieval precision. Finding the sweet spot saves money.

The Chunking Decision Tree

Your choice of chunking strategy depends on your document type and use case:

Document Type?
Decide based on structure
Plain text โ†™ Code โ†“ HTML/Markdown โ†˜ PDF
Semantics Important?
Is topic coherence critical?
No โ†’ Fixed-Size | Yes โ†’ Semantic/Recursive
Latency/Cost?
Can you afford embeddings?
Speed Critical โ†’ Recursive | High Budget โ†’ Semantic
Final Strategy
Implement & optimize

Key Insight: There is no universal “best” chunking strategy. The optimal approach depends on your documents, your queries, your LLM, and your retrieval model. The solution is: measure, experiment, iterate.

Core Concepts and Terminology

Before diving into specific strategies, understand the key concepts that define chunking behavior.

1. Chunk Size

The number of tokens (or characters) in each chunk. Typical range: 128-2048 tokens.

Small Chunks (128-256 tokens)

Pros: High retrieval precision, less irrelevant context. Cons: May lack context, requires more embeddings, higher query latency.

Medium Chunks (512 tokens)

Pros: Balance of context and precision, industry standard. Cons: May miss long-range dependencies.

Large Chunks (1024+ tokens)

Pros: Rich context, fewer embeddings. Cons: Lower precision, may include irrelevant content, slower retrieval.

2. Chunk Overlap

The number of tokens repeated at the boundary between consecutive chunks. Preserves context and maintains continuity.

Overlap Calculation
overlap_tokens = chunk_size * overlap_percent Example: 512-token chunks with 25% overlap = 128 tokens of overlap Chunk 1: tokens 0-511 Chunk 2: tokens 384-895 (overlap from 384-511) Chunk 3: tokens 768-1279 (overlap from 768-895)

Why Overlap Matters

Overlap ensures that important information split across chunk boundaries doesn't get lost. Without overlap, a sentence split between chunk N and chunk N+1 might only appear in one chunk's embedding. With overlap, it appears in both, improving retrieval probability.

3. Semantic Coherence

A good chunk contains related, cohesive information. A bad chunk arbitrarily splits meaning.

Bad Split Example:

What NOT to do
Chunk 1: "The transformer architecture introduced in 2017 has become the foundation of modern... [cuts at 512 tokens]" Chunk 2: "...NLP systems. Self-attention enables parallel processing. Unlike RNNs which process sequentially, transformers can process entire..." [continues]

Good Split Example:

What to do
Chunk 1: "The transformer architecture, introduced in 2017 by Vaswani et al., replaced RNNs as the dominant sequence model. Self-attention enables parallel processing, making transformers dramatically faster to train." Chunk 2: "Unlike RNNs which process tokens sequentially, transformers process entire sequences simultaneously. This parallelization, combined with layer normalization and residual connections, enables training on massive datasets..."

4. Metadata Preservation

Preserving source document metadata (title, section, URL, author) enables:

  • Better answer attribution ("According to Section 4.2...")
  • Filtering retrieved chunks by document source
  • Tracing retrieval decisions back to source
  • Evaluating which documents are most helpful

5. The Retrieval Quality Spectrum

Irrelevant (0-20%)
Retrieved content doesn't answer the query
Partially Relevant (20-60%)
Contains some useful info but mixed with irrelevant content
Relevant (60-85%)
Mostly relevant with some tangential content
Highly Relevant (85-100%)
Directly answers the query with minimal noise

The Chunking Challenge: You're optimizing for retrieval quality โ€” finding chunks that are relevant to a query โ€” but you won't know query intent until runtime. This is why choosing the right chunking strategy upfront is so important.

Fixed-Size Chunking with Overlap

The simplest chunking strategy: split documents into chunks of fixed size (e.g., 512 tokens) with optional overlap (e.g., 20% overlap = 102 tokens).

How It Works

Iterate through the document, creating chunks at regular intervals. When using overlap, the end of one chunk repeats at the start of the next.

Python โ€” Fixed-Size Chunking Implementation
def fixed_size_chunking(text, chunk_size=512, overlap=0.2): """ Split text into fixed-size chunks with overlap. Args: text: Input document text chunk_size: Size of each chunk in tokens (approximate) overlap: Overlap ratio (0.0 to 1.0) Returns: List of chunk strings """ # Split into words (approximation of tokens) words = text.split() overlap_size = int(chunk_size * overlap) step_size = chunk_size - overlap_size chunks = [] for i in range(0, len(words), step_size): chunk = words[i:i + chunk_size] chunks.append(" ".join(chunk)) if i + chunk_size >= len(words): break return chunks # Example usage document = """ The transformer architecture revolutionized NLP. Self-attention mechanisms enable parallel processing... [... more text ...] """ chunks = fixed_size_chunking(document, chunk_size=128, overlap=0.2) for i, chunk in enumerate(chunks): print(f"Chunk {i}: {len(chunk.split())} words") print(f" {chunk[:80]}...\n")

Advantages

  • Simplicity: Easy to implement and understand
  • Speed: No computation overhead, runs instantly
  • Predictability: Chunk count is predictable (useful for cost estimation)
  • Flexibility: Works with any document type

Disadvantages

  • Semantic Ignorance: Doesn't respect paragraph/section boundaries
  • Context Loss: May split sentences or ideas mid-way
  • One-Size-Fits-All: Same strategy regardless of document structure
  • Lower Retrieval Quality: Typically 60-70% retrieval accuracy

When to Use

Fixed-size chunking is appropriate when:

  • You're prototyping quickly and don't need optimal quality
  • Documents are highly homogeneous plain text
  • You need to be language-agnostic (no tokenizer dependencies)
  • You're optimizing purely for speed over quality

Exercise: Fixed-Size with Different Overlap

Implement fixed-size chunking with 0%, 20%, and 50% overlap on a sample document. Compare the number of chunks and overlap between consecutive chunks.

Python โ€” Starter Code
document = "Your text here..." for overlap_pct in [0, 0.2, 0.5]: chunks = fixed_size_chunking(document, 256, overlap_pct) print(f"Overlap {overlap_pct*100}%: {len(chunks)} chunks")

Token Counting

The example uses word count as an approximation of token count. In production, use a proper tokenizer (tiktoken for GPT, sentencepiece for others) to match your embedding model's tokenization.

Recursive Chunking

Recursive chunking respects document structure by attempting to split on progressively more granular separators (paragraph โ†’ sentence โ†’ word โ†’ character) until chunks fit the target size.

How It Works

The algorithm maintains a list of separator patterns (hierarchical). It tries the first separator; if chunks are still too large, it tries the next separator within those chunks, recursively, until all chunks fit the target size.

Python โ€” Recursive Chunking with LangChain
from langchain.text_splitter import RecursiveCharacterTextSplitter # Create a recursive splitter splitter = RecursiveCharacterTextSplitter( chunk_size=512, chunk_overlap=102, # ~20% overlap separators=[ "\n\n", # Paragraph "\n", # Line break ". ", # Sentence " ", # Word "" # Character ] ) document = """ Chapter 1: Introduction Paragraph 1 about transformers... Paragraph 2 about attention... """ chunks = splitter.split_text(document) for i, chunk in enumerate(chunks): print(f"Chunk {i} ({len(chunk)} chars): {chunk[:60]}...")

Why Recursive Splitting?

Semantic Preservation

By preferring paragraph splits over sentence splits, recursive splitting keeps related paragraphs together when possible.

Hierarchy Awareness

The separator list encodes document structure: chapters > paragraphs > sentences > words. Respects this hierarchy when possible.

Flexibility

Automatically adapts. If paragraphs are too large, it splits sentences. If sentences are still too large, it splits words.

Separator Strategies for Different Document Types

Separators by Document Type
# Markdown & Documentation markdown_seps = ["\n\n", "\n", "# ", "## ", "### ", ". ", " ", ""] # Code Files code_seps = ["\n\n", "\nclass ", "\ndef ", "\n\n", "\n", " ", ""] # HTML/XML html_seps = ["\n\n", "\n", "
", "

", "
", ". ", " ", ""] # Plain Text (Default) text_seps = ["\n\n", "\n", ". ", " ", ""]

Advantages

Disadvantages

Exercise: Recursive Chunking

Use RecursiveCharacterTextSplitter with a markdown document. Try different separator lists and observe how chunk boundaries change.

Python โ€” Starter Code
from langchain.text_splitter import RecursiveCharacterTextSplitter markdown_text = ''' # Chapter 1 ## Section 1.1 Content here... ''' splitter = RecursiveCharacterTextSplitter(chunk_size=200) chunks = splitter.split_text(markdown_text)

Semantic Chunking with Embeddings

Semantic chunking uses embeddings to measure semantic similarity between sentences. Chunks are boundaries where semantic similarity drops below a threshold, preserving semantic coherence.

How It Works

  1. Split document into sentences
  2. Generate embeddings for each sentence
  3. Compute cosine similarity between consecutive sentences
  4. Identify breakpoints where similarity drops below threshold
  5. Group sentences into chunks based on these breakpoints
Python โ€” Semantic Chunking Implementation
import numpy as np from sklearn.metrics.pairwise import cosine_similarity import re def semantic_chunking(text, embedding_func, threshold=0.5, max_size=512): """ Chunk text using semantic similarity. Args: text: Input document embedding_func: Function that returns embeddings (e.g., from OpenAI) threshold: Similarity threshold for chunk boundaries (0-1) max_size: Maximum tokens per chunk Returns: List of chunks """ # Split into sentences sentences = re.split(r'(?<=[.!?])\s+', text) sentences = [s.strip() for s in sentences if s.strip()] if len(sentences) == 0: return [] # Get embeddings for each sentence embeddings = [] for sentence in sentences: emb = embedding_func(sentence) embeddings.append(emb) # Compute similarities between consecutive sentences similarities = [] for i in range(len(embeddings) - 1): sim = cosine_similarity( [embeddings[i]], [embeddings[i+1]] )[0][0] similarities.append(sim) # Identify chunk boundaries where similarity drops breakpoints = [0] for i, sim in enumerate(similarities): if sim < threshold: breakpoints.append(i + 1) breakpoints.append(len(sentences)) # Create chunks from sentences between breakpoints chunks = [] for start, end in zip(breakpoints[:-1], breakpoints[1:]): chunk = " ".join(sentences[start:end]) if len(chunk) > max_size: # If still too large, split recursively chunks.extend(semantic_chunking( chunk, embedding_func, threshold, max_size )) else: chunks.append(chunk) return chunks

Example with OpenAI Embeddings

Python โ€” Semantic Chunking with OpenAI
from openai import OpenAI client = OpenAI() def get_embedding(text): response = client.embeddings.create( input=text, model="text-embedding-3-small" ) return response.data[0].embedding document = """ The transformer architecture has revolutionized AI. It introduced self-attention mechanisms that enable parallel processing. Unlike RNNs, transformers process entire sequences at once. This architectural choice, combined with scale, unlocked the modern AI era. """ chunks = semantic_chunking(document, get_embedding, threshold=0.6) for i, chunk in enumerate(chunks): print(f"Chunk {i}: {chunk[:70]}...")

Advantages

Disadvantages

Threshold Tuning

The similarity threshold dramatically affects chunk boundaries:

Threshold = 0.3
Many small chunks, low latency
Threshold = 0.5
Balanced chunks, balanced latency
Threshold = 0.8
Few large chunks, high latency

Exercise: Semantic Chunking

Implement semantic chunking with different thresholds (0.3, 0.5, 0.8). Compare the number of chunks and their sizes.

Python โ€” Starter Code
# Use mock embeddings for testing import numpy as np def mock_embedding(text): np.random.seed(hash(text) % 2**32) return np.random.randn(384)

Document-Aware Chunking

Different document types (HTML, PDF, code, markdown) have distinct structure. Document-aware chunking preserves this structure, respecting hierarchy, formatting, and semantic units specific to each type.

HTML Document Chunking

Python โ€” HTML-Aware Chunking
from bs4 import BeautifulSoup from langchain.document_loaders import UnstructuredHTMLLoader html_content = """ <article> <h1>Machine Learning Fundamentals</h1> <section id="intro-2"> <h2>Introduction</h2> <p>Machine learning is...</p> <p>There are three main types...</p> </section> <section id="supervised"> <h2>Supervised Learning</h2> <p>In supervised learning...</p> </section> </article> """ def chunk_html(html_text, chunk_size=512): soup = BeautifulSoup(html_text, 'html.parser') chunks = [] # Extract sections preserving hierarchy for section in soup.find_all(['section', 'article']): heading = section.find(['h1', 'h2', 'h3']) heading_text = heading.text if heading else "" paragraphs = section.find_all('p') content = heading_text + "\n" + "\n".join(p.text for p in paragraphs) # Further chunk if too large if len(content) > chunk_size: words = content.split() for i in range(0, len(words), chunk_size//10): chunks.append(" ".join(words[i:i+chunk_size//10])) else: chunks.append(content) return chunks chunks = chunk_html(html_content) for i, chunk in enumerate(chunks): print(f"Chunk {i}: {chunk[:80]}...")

Code Document Chunking

Python โ€” Code-Aware Chunking
import re def chunk_code(code_text, chunk_size=512): """ Chunk code while preserving class and function boundaries. """ chunks = [] current_chunk = "" # Split on class/function definitions blocks = re.split(r'^(class |def )', code_text, flags=re.MULTILINE) for block in blocks: if len(current_chunk) + len(block) <= chunk_size: current_chunk += block else: if current_chunk: chunks.append(current_chunk) current_chunk = block if current_chunk: chunks.append(current_chunk) return chunks code = """ class DataProcessor: def __init__(self): self.data = [] def process(self, item): return item.upper() def helper_function(): return "result" """ chunks = chunk_code(code, chunk_size=100) for i, chunk in enumerate(chunks): print(f"Code Chunk {i}:\n{chunk}\n---")

PDF Document Chunking

Python โ€” PDF-Aware Chunking
from PyPDF2 import PdfReader from langchain.document_loaders import PyPDFLoader def chunk_pdf(pdf_path, chunk_size=512): """ Extract text from PDF preserving page and section structure. """ loader = PyPDFLoader(pdf_path) documents = loader.load() chunks = [] for doc in documents: text = doc.page_content metadata = doc.metadata # Includes page number # Chunk while preserving metadata words = text.split() for i in range(0, len(words), chunk_size): chunk_text = " ".join(words[i:i+chunk_size]) chunks.append({ "text": chunk_text, "page": metadata.get("page", 0), "source": metadata.get("source", "") }) return chunks # Usage: # chunks = chunk_pdf("document.pdf", chunk_size=512)

Metadata Enrichment

Document-aware chunking enables metadata enrichment, improving attribution and filtering:

Python โ€” Chunk with Rich Metadata
chunk_with_metadata = { "text": "The transformer architecture...", "chunk_id": "doc_001_chunk_005", "document": "learn-transformers.html", "section": "Core Concepts", "subsection": "Self-Attention", "page": 3, "start_char": 1250, "end_char": 1762, "language": "en", "document_type": "html" }

Advantages

Disadvantages

Exercise: Multi-Format Chunking

Implement a chunking pipeline that handles markdown, HTML, and plain text, preserving structure for each format.

Python โ€” Starter Code
def universal_chunk(text, format_type, chunk_size=512): if format_type == 'html': return chunk_html(text, chunk_size) elif format_type == 'markdown': return chunk_markdown(text, chunk_size) else: return fixed_size_chunking(text, chunk_size)

Agentic Chunking

Agentic chunking uses an LLM as an agent to intelligently parse documents, deciding how to chunk based on semantic meaning rather than mechanical rules.

How It Works

The LLM reads a document and decides:

  1. Which content belongs together
  2. Where natural boundaries exist
  3. Which details are important vs contextual
  4. How to create chunks with necessary context
Python โ€” Agentic Chunking with Claude
from anthropic import Anthropic client = Anthropic() def agentic_chunking(document_text, model="claude-3-5-sonnet-20241022"): """ Use Claude to intelligently chunk a document. """ response = client.messages.create( model=model, max_tokens=4096, messages=[ { "role": "user", "content": f"""Analyze this document and identify natural semantic chunks. For each chunk: 1. Define clear boundaries 2. Identify the main topic 3. List key concepts 4. Note any required context Document: {document_text} Provide your response as structured JSON with chunks array.""" } ] ) # Parse LLM response to extract chunks response_text = response.content[0].text # Extract JSON from response import json import re json_match = re.search(r'\{.*\}', response_text, re.DOTALL) if json_match: chunks_data = json.loads(json_match.group()) return chunks_data["chunks"] return [] document = """ Chapter 1: Neural Networks A neural network is a computational model... Neurons process information... Chapter 2: Training To train a neural network, we use gradient descent... Backpropagation computes gradients... """ chunks = agentic_chunking(document) for i, chunk in enumerate(chunks): print(f"Chunk {i}: {chunk.get('topic', 'Unknown')}") print(f" Content: {chunk.get('text', '')[:100]}...")

Advanced: Multi-Pass Agentic Chunking

Python โ€” Multi-Pass Analysis
def multi_pass_agentic_chunking(document, max_passes=3): """ Iteratively refine chunking through multiple LLM passes. Pass 1: Identify high-level structure Pass 2: Refine boundaries based on semantic analysis Pass 3: Verify chunk quality and optimize """ chunks = [] for pass_num in range(max_passes): if pass_num == 0: prompt = f"""Identify the main topics and high-level structure. Document: {document}""" elif pass_num == 1: prompt = f"""Refine chunk boundaries to be semantically coherent. Previous analysis: {chunks}""" else: prompt = f"""Verify chunk quality. Are chunks: 1. Semantically coherent? 2. Self-contained? 3. Complete (not mid-sentence)? Current chunks: {chunks}""" response = client.messages.create( model="claude-3-5-sonnet-20241022", max_tokens=2048, messages=[{"role": "user", "content": prompt}] ) # Update chunks based on LLM feedback # (simplified; real implementation would parse structured output) return chunks

Advantages

Disadvantages

Cost Analysis

Fixed-Size
$0.00 (free)
Semantic (embeddings)
$0.02-0.05 per 100K tokens
Agentic (Claude)
$0.30-1.00 per 100K tokens

Approximate cost for processing 100K tokens

When to use Agentic Chunking: When document quality is critical (legal documents, medical records), or when documents are complex and standard strategies fail. Not suitable for high-volume processing with tight budgets.

Chunking Strategy Comparison

Here's a comprehensive comparison of all five strategies across key dimensions:

Dimension Fixed-Size Recursive Semantic Doc-Aware Agentic
Retrieval Accuracy 62% 78% 91% 85% 96%
Speed Instant (ms) Fast (ms) Moderate (secs) Moderate (secs) Slow (mins)
Cost Free Free $0.02/100K $0.02/100K $0.50/100K
Implementation Complexity Very Easy Easy Moderate Moderate Complex
Deterministic Yes Yes Yes Yes No (LLM variance)
Best For Prototyping, simple documents Structured text, markdown Mixed documents, high quality Specific formats (HTML, PDF) Critical documents, best quality

Decision Matrix

Use this matrix to choose the right strategy for your use case:

Choose Fixed-Size if:
  • Building a rapid prototype
  • Need zero overhead
  • Documents are homogeneous
Choose Recursive if:
  • Documents are structured (markdown, code)
  • Need decent quality without cost
  • Want simple implementation
Choose Semantic if:
  • High retrieval quality needed
  • Have embedding budget
  • Mixed document types
Choose Agentic if:
  • Critical documents (legal, medical)
  • Quality > cost/latency
  • Complex domain knowledge needed

Chunk Size Optimization Experiment

The optimal chunk size is not universal โ€” it depends on your documents, queries, retrieval model, and LLM. The best approach is empirical: test different sizes and measure actual retrieval quality.

Experiment Setup

Here's a complete framework for optimizing chunk size:

Python โ€” Chunk Size Optimization Framework
import json import time from sklearn.metrics import ndcg_score from tqdm import tqdm def run_chunk_optimization(documents, queries, chunk_sizes, embedding_func): """ Test different chunk sizes and measure retrieval quality. Returns: Dictionary of results by chunk size """ results = {} for chunk_size in chunk_sizes: print(f"\nTesting chunk_size={chunk_size}") start_time = time.time() # Create chunks all_chunks = [] for doc in documents: chunks = semantic_chunking(doc, embedding_func, max_size=chunk_size) all_chunks.extend(chunks) chunking_time = time.time() - start_time # Embed all chunks embeddings = [] for chunk in all_chunks: emb = embedding_func(chunk) embeddings.append(emb) # Evaluate on queries ndcg_scores = [] for query in tqdm(queries, desc=f"Evaluating {len(queries)} queries"): query_emb = embedding_func(query["text"]) # Find top-k retrieved chunks similarities = [cosine_similarity([query_emb], [emb])[0][0] for emb in embeddings] ranked_indices = sorted(range(len(similarities)), key=lambda i: similarities[i], reverse=True)[:5] # Compute NDCG relevances = [1 if idx in query["relevant_chunk_ids"] else 0 for idx in ranked_indices] ndcg = ndcg_score([query["relevances"]], [relevances]) ndcg_scores.append(ndcg) results[chunk_size] = { "avg_ndcg": sum(ndcg_scores) / len(ndcg_scores), "num_chunks": len(all_chunks), "chunking_time": chunking_time, "chunks_per_query": len(all_chunks) / len(queries) } return results # Run optimization chunk_sizes = [128, 256, 512, 1024, 2048] results = run_chunk_optimization(documents, test_queries, chunk_sizes, get_embedding) # Print results print("\n=== Optimization Results ===") for size in chunk_sizes: r = results[size] print(f"Size {size:4d}: NDCG={r['avg_ndcg']:.3f}, " f"Chunks={r['num_chunks']:4d}, " f"Time={r['chunking_time']:.2f}s")

Key Metrics to Track

NDCG@k Score

Normalized Discounted Cumulative Gain measures ranking quality. Score from 0-1, higher is better. Standard in IR evaluation.

MRR (Mean Reciprocal Rank)

Average position of the first relevant chunk. Measures how quickly you find relevant results.

Recall@k

Percentage of relevant chunks in top-k results. Critical for RAG (need relevant info to generate good answers).

Latency

Time to chunk documents and query the index. Includes embedding time and database queries.

Typical Results from Optimization

128 tokens
NDCG: 0.82
256 tokens
NDCG: 0.87
512 tokens
NDCG: 0.89
1024 tokens
NDCG: 0.86
2048 tokens
NDCG: 0.78

Typical NDCG@5 scores by chunk size (results vary by domain)

Optimization Insights

Sweet Spot

In most cases, 256-512 tokens yields the best NDCG scores. Smaller chunks miss context; larger chunks include noise. Your optimization should confirm this for your specific data.

Domain Sensitivity

Technical documents benefit from larger chunks (512+) to maintain context. News articles and short documents benefit from smaller chunks (128-256). Always test on representative data from your domain.

Overlap Impact

Overlap improves recall at the cost of more total chunks. 20-25% overlap typically provides good balance. Too much overlap (>50%) creates excessive redundancy without quality gains.

Retrieval Quality by Chunk Size

This section visualizes how chunk size affects the three critical dimensions of retrieval quality:

1. Precision vs Recall Tradeoff

Precision by Chunk Size

128 tokens
92% Precision@5
256 tokens
88% Precision@5
512 tokens
84% Precision@5
1024 tokens
76% Precision@5

Recall by Chunk Size

128 tokens
65% Recall@5
256 tokens
78% Recall@5
512 tokens
86% Recall@5
1024 tokens
91% Recall@5

2. Latency by Chunk Size and Strategy

Fixed-Size (512 tokens)
2ms
Recursive (512 tokens)
5ms
Semantic (512 tokens)
18ms (embeddings)
Agentic (512 tokens)
60s+ (LLM latency)

3. Cost Analysis: Chunking + Embedding + Storage

Scenario: Process 1 million documents (100 tokens each) and serve 1000 queries/day

Fixed-Size

Chunking: Free

Embedding: N/A

Storage: 1M chunks

Total: $0

Semantic

Chunking: Free

Embedding: 100M tokens โ†’ $20

Storage: 1M vectors

Total: $20 (+$5/mo)

Agentic

Chunking: 100M tokens โ†’ $300

Embedding: 1M chunks โ†’ $20

Storage: 1M vectors

Total: $320

Advanced Chunking Techniques

1. Sliding Window Chunking

Maximize context by overlapping chunks more intelligently. Instead of uniform overlap, use sliding windows with variable step sizes.

Python โ€” Sliding Window Chunking
def sliding_window_chunking(text, window_size=512, step_size=256): """ Create overlapping chunks using sliding window approach. Args: text: Input document window_size: Size of each window step_size: How far to move between windows Returns: List of chunks with indices """ words = text.split() chunks_with_info = [] for start in range(0, len(words) - window_size, step_size): chunk = " ".join(words[start:start + window_size]) chunks_with_info.append({ "text": chunk, "start_idx": start, "end_idx": start + window_size, "overlap_with_next": window_size - step_size }) return chunks_with_info

2. Hierarchical Chunking

Create a hierarchy of chunks: summary chunks, section chunks, paragraph chunks. At query time, start with summary chunks to narrow scope, then retrieve detailed chunks.

Python โ€” Hierarchical Chunking
def hierarchical_chunking(document): """ Create multi-level chunk hierarchy. """ hierarchy = { "document_summary": summarize(document), "sections": [], } for section in document.sections: section_data = { "title": section.title, "summary": summarize(section.text), "paragraphs": [] } for paragraph in section.paragraphs: section_data["paragraphs"].append({ "text": paragraph, "summary": summarize(paragraph) }) hierarchy["sections"].append(section_data) return hierarchy # Usage at retrieval time: # 1. Search document summaries # 2. Search matching section summaries # 3. Search paragraph chunks

3. Colbert Chunking (Contextual Dense Embedding)

ColBERT uses late interaction for dense retrieval. Instead of embedding entire chunks, embed individual tokens and compute interaction at query time. Enables finer-grained chunking without quality loss.

4. Graph-Based Chunking

Represent documents as knowledge graphs and chunk based on semantic graph structure rather than linear text order.

Node Parsers and LlamaIndex Integration

LlamaIndex provides production-grade node parsers that abstract chunking complexity. Nodes are the LlamaIndex term for chunks โ€” they include the text, metadata, and embedding vectors.

Simple Node Parser

Python โ€” LlamaIndex SimpleNodeParser
from llama_index.core.node_parser import SimpleNodeParser # Create a simple parser (fixed-size chunking) parser = SimpleNodeParser.from_defaults( chunk_size=512, chunk_overlap=102 ) # Parse documents into nodes from llama_index.core import Document docs = [ Document(text="Your document text here..."), Document(text="Another document...") ] nodes = parser.get_nodes_from_documents(docs) for node in nodes: print(f"Node: {node.node_id}") print(f"Text: {node.text[:100]}...") print(f"Metadata: {node.metadata}") print()

Semantic Chunking with LlamaIndex

Python โ€” SemanticSplitterNodeParser
from llama_index.core.node_parser import SemanticSplitterNodeParser from llama_index.embeddings.openai import OpenAIEmbedding # Create semantic splitter embed_model = OpenAIEmbedding(model="text-embedding-3-small") parser = SemanticSplitterNodeParser( buffer_size=1, breakpoint_percentile_threshold=95, embed_model=embed_model, ) nodes = parser.get_nodes_from_documents(docs) # Each node has: # - node_id: Unique identifier # - text: The chunk content # - metadata: Including source document, page, etc. # - embedding: The vector embedding

Hierarchical Node Parser

Python โ€” HierarchicalNodeParser
from llama_index.core.node_parser import HierarchicalNodeParser # Create hierarchical parser # Creates summary nodes at different levels parser = HierarchicalNodeParser.from_defaults( chunk_size=512, chunk_overlap=102, ) nodes = parser.get_nodes_from_documents(docs) # Includes both leaf nodes (actual text) and parent nodes (summaries) for node in nodes: if hasattr(node, 'is_parent') and node.is_parent: print(f"Summary Node: {node.text[:80]}...") else: print(f"Leaf Node: {node.text[:80]}...")

Integration with RAG Pipeline

Python โ€” Complete RAG with Proper Chunking
from llama_index.core import VectorStoreIndex # Parse documents parser = SemanticSplitterNodeParser( buffer_size=1, breakpoint_percentile_threshold=95, embed_model=embed_model, ) nodes = parser.get_nodes_from_documents(docs) # Create index from parsed nodes index = VectorStoreIndex(nodes) # Query query_engine = index.as_query_engine() response = query_engine.query("How do transformers work?") print(response)

Node Properties and Metadata

Node Structure
{ "node_id": "abc123", "text": "The chunk content...", "metadata": { "page_label": "1", "file_name": "document.pdf", "doc_id": "doc_001", "section": "Introduction", "start_char_idx": 0, "end_char_idx": 512 }, "relationships": { "parent": "parent_node_id", "next": "next_node_id", "prev": "prev_node_id" }, "embedding": [0.1, 0.2, ..., 0.9] }

Custom Node Parser

Python โ€” Creating Custom Parsers
from llama_index.core.node_parser import NodeParser from llama_index.core.schema import BaseNode class CustomNodeParser(NodeParser): def _parse_nodes(self, docs, **kwargs): """Implement your custom chunking logic.""" nodes = [] for doc in docs: # Your custom chunking here chunks = your_chunking_function(doc.text) for i, chunk_text in enumerate(chunks): node = BaseNode( text=chunk_text, metadata={ "chunk_idx": i, "doc_id": doc.doc_id, **doc.metadata } ) nodes.append(node) return nodes # Use your parser parser = CustomNodeParser() nodes = parser.get_nodes_from_documents(docs)

Best Practices for Production Chunking

1. Always Test on Representative Data

Create a test set of documents and queries that reflect your actual use case. Chunking decisions must be validated on real data, not synthetic examples.

Best Practice

Use your production queries to evaluate chunking quality. Run A/B tests with different chunk sizes and strategies. Track NDCG, MRR, and Recall@5. Choose the strategy with best overall quality/cost tradeoff.

2. Preserve Document Structure and Metadata

Python โ€” Chunk with Rich Metadata
chunk = { "text": "The chunk content...", "source_doc": "technical_guide.pdf", "page": 5, "section": "3.2 Architecture", "timestamp": "2024-02-27", "language": "en", "quality_score": 0.92 }

3. Implement Versioning for Chunks

When you update chunking strategy, version your chunks. This allows rolling back if a new strategy performs worse.

Python โ€” Chunk Versioning
chunk = { "chunk_id": "v2_doc_001_chunk_005", "version": 2, # Incremented when strategy changes "strategy": "semantic-0.6-threshold", "text": "...", "created_at": "2024-02-27T10:30:00Z" } # At query time, can select specific versions: # SELECT * FROM chunks WHERE version = 2

4. Monitor Chunk Quality Metrics

Track these metrics continuously in production:

Retrieval Quality

NDCG@5, MRR, Recall@5. Should remain consistent over time.

Chunk Size Distribution

Mean, median, p95 chunk sizes. Should match your target.

Answer Quality

BLEU, ROUGE scores of final LLM answers vs gold answers.

Latency

End-to-end latency from query to retrieved chunks. Should be <100ms.

5. Handle Edge Cases

Common Edge Cases
# Very long documents (>10K tokens) - Implement hierarchical chunking - Or split into subdocuments first # Very short documents (<128 tokens) - Don't chunk, use entire document as one chunk - Or add padding/context from related documents # Tables and structured data - Parse tables separately, include col headers in chunks - Don't split rows arbitrarily # Code with complex nesting - Respect bracket/indentation structure - Include full function/class definitions, not partial # Multiple languages - Use language-specific tokenizers - May need different chunk sizes per language

6. Optimize for Your Retrieval Model

The best chunk size depends on your retrieval model:

7. Implement Chunk Quality Filtering

Python โ€” Filter Low-Quality Chunks
def should_keep_chunk(chunk_text, min_words=5, max_redundancy=0.3): """Filter out chunks that don't meet quality thresholds.""" # Too short if len(chunk_text.split()) < min_words: return False # Too much whitespace if len(chunk_text.strip()) < len(chunk_text) * 0.5: return False # Mostly punctuation/numbers alphanumeric = sum(1 for c in chunk_text if c.isalnum()) if alphanumeric < len(chunk_text) * 0.3: return False # Check for repeated sentences (redundancy) sentences = chunk_text.split(". ") if len(sentences) > 0: unique_ratio = len(set(sentences)) / len(sentences) if unique_ratio < (1 - max_redundancy): return False return True # Filter chunks before indexing quality_chunks = [c for c in chunks if should_keep_chunk(c)]

Common Mistakes and How to Avoid Them

Mistake 1: Using Fixed-Size Chunking for Everything

Problem: Fixed-size ignores document structure. A 512-token chunk may cut paragraphs and concepts mid-way.

Solution: Use recursive or semantic chunking for heterogeneous documents. Fixed-size only for homogeneous plain text.

Mistake 2: No Overlap Between Chunks

Problem: Information split between chunks may be missed. Queries aligned with chunk boundaries may lose relevant context.

Solution: Always use 20-25% overlap. The ~2x cost in total chunks is worth the retrieval quality improvement.

Mistake 3: Choosing Chunk Size Without Testing

Problem: Arbitrary choices (e.g., "use 1024 tokens") ignore your specific data and use case.

Solution: Run optimization experiments as shown in Section 10. Test 128, 256, 512, 1024 on representative data.

Mistake 4: Losing Source Attribution

Problem: If chunks don't preserve document/page metadata, you can't cite sources in answers.

Solution: Always include source_doc, page, section, and character offsets in chunk metadata. Pass this through to final answers.

Mistake 5: Not Handling Document-Specific Formats

Problem: Using generic chunking on HTML, PDF, or code loses structure. HTML tags get included in embeddings. PDF page breaks aren't preserved.

Solution: Use document-aware parsers. Use LangChain loaders and LlamaIndex node parsers that handle format-specific structure.

Mistake 6: Using Inconsistent Tokenization

Problem: Word count โ‰  token count. If you chunk by word count but your embedding model uses byte-pair encoding (BPE), actual token count may be 1.3x higher.

Solution: Use the same tokenizer as your embedding model. Use tiktoken for OpenAI models, use HF transformers tokenizers for others.

Mistake 7: Semantic Chunking Without Threshold Tuning

Problem: The default threshold (0.5) may not be optimal for your documents. Using it without tuning wastes the semantic approach's potential.

Solution: Analyze similarity distributions in your documents. Test thresholds from 0.3 to 0.8. Pick the one with best NDCG on test queries.

Mistake 8: Forgetting to Re-Chunk When Documents Update

Problem: If a document is updated but chunks aren't regenerated, your index becomes stale and questions about updated content get wrong answers.

Solution: Implement document versioning and incremental chunking. Detect updated documents and re-chunk only those. Invalidate old embeddings.

Complete Code Examples and Recipes

Recipe 1: End-to-End RAG with Optimal Chunking

Python โ€” Production RAG Pipeline
from llama_index.core import VectorStoreIndex, Document from llama_index.core.node_parser import SemanticSplitterNodeParser from llama_index.embeddings.openai import OpenAIEmbedding from llama_index.llms.openai import OpenAI import os # Initialize models embed_model = OpenAIEmbedding(model="text-embedding-3-small") llm = OpenAI(model="gpt-4-turbo", temperature=0.7) # Load documents documents = [ Document(text="Your document content...", metadata={"source": "file1.pdf", "page": 1}), Document(text="Another document...", metadata={"source": "file2.pdf", "page": 1}), ] # Create semantic chunking parser parser = SemanticSplitterNodeParser( buffer_size=1, breakpoint_percentile_threshold=95, embed_model=embed_model, ) # Parse documents into nodes nodes = parser.get_nodes_from_documents(documents) # Create vector index index = VectorStoreIndex(nodes) # Query query_engine = index.as_query_engine( llm=llm, similarity_top_k=5, ) response = query_engine.query("What is semantic chunking?") print(response)

Recipe 2: Chunking with Quality Filtering

Python โ€” Quality-Aware Chunking
def chunk_with_quality_filter(text, chunk_size=512, overlap=0.2, min_quality=0.7): """Chunk text and filter by quality.""" # Initial chunking chunks = semantic_chunking(text, chunk_size=chunk_size, overlap=overlap) # Compute quality scores quality_chunks = [] for chunk in chunks: score = compute_quality_score(chunk) if score >= min_quality: quality_chunks.append({ "text": chunk, "quality_score": score, "length": len(chunk.split()) }) return quality_chunks def compute_quality_score(chunk): """Estimate chunk quality (0-1).""" score = 0 # Reward completeness if chunk.endswith((".","!","?")): score += 0.2 # Reward coherence (sentence count) sentences = len(chunk.split(".")) if 2 <= sentences <= 5: score += 0.3 # Reward information density unique_words = len(set(chunk.lower().split())) total_words = len(chunk.split()) if unique_words / total_words > 0.6: score += 0.5 return min(score, 1.0)

Recipe 3: Multi-Format Document Processing

Python โ€” Universal Document Chunker
from pathlib import Path from PyPDF2 import PdfReader from bs4 import BeautifulSoup def chunk_any_document(file_path, chunk_size=512): """Intelligently chunk any document type.""" path = Path(file_path) if path.suffix == ".pdf": text = extract_text_from_pdf(file_path) elif path.suffix == ".html": text = extract_text_from_html(file_path) elif path.suffix == ".md": text = extract_text_from_markdown(file_path) else: text = path.read_text() # Use semantic chunking for all chunks = semantic_chunking(text, chunk_size=chunk_size) # Preserve source metadata return [ {"text": chunk, "source": str(file_path), "format": path.suffix} for chunk in chunks ] def extract_text_from_pdf(pdf_path): reader = PdfReader(pdf_path) text = "" for page in reader.pages: text += page.extract_text() return text def extract_text_from_html(html_path): soup = BeautifulSoup(open(html_path).read(), "html.parser") # Remove script/style tags for tag in soup(["script", "style"]): tag.decompose() return soup.get_text() def extract_text_from_markdown(md_path): # Simple approach: just read text # More robust: use markdown parser to preserve structure return open(md_path).read()

Hands-On Exercises

Exercise 1: Implement Fixed-Size Chunking

Goal: Build a fixed-size chunker with overlap from scratch.

Fixed-Size Implementation

Implement fixed_size_chunking() with parameters for chunk_size and overlap. Test with different overlap values (0%, 20%, 50%) and verify the overlapping content.

Python โ€” Starter Code
def fixed_size_chunking(text, chunk_size=512, overlap=0.2): words = text.split() overlap_size = int(chunk_size * overlap) step_size = chunk_size - overlap_size chunks = [] # TODO: Implement your logic here return chunks # Test doc = "The transformer architecture...\n[100+ words]" chunks = fixed_size_chunking(doc, 256, 0.2) print(f"Created {len(chunks)} chunks") # Verify overlap if len(chunks) > 1: # Check that chunk[0] ends with content from chunk[1] pass

Exercise 2: Semantic Chunking with Threshold Tuning

Goal: Implement semantic chunking and find optimal threshold for a document.

Semantic Threshold Optimization

Implement semantic_chunking() with variable threshold. Test thresholds 0.3, 0.5, 0.7, 0.9 and plot: number of chunks vs. threshold.

Python โ€” Starter Code
from sklearn.metrics.pairwise import cosine_similarity def semantic_chunking(text, embedding_func, threshold=0.5): # TODO: Implement semantic chunking # 1. Split into sentences # 2. Get embeddings # 3. Compute similarities # 4. Break where similarity < threshold pass # Experiment thresholds = [0.3, 0.5, 0.7, 0.9] chunk_counts = [] for threshold in thresholds: chunks = semantic_chunking(document, get_embedding, threshold) chunk_counts.append(len(chunks)) print(f"Threshold {threshold}: {len(chunks)} chunks") # Plot: threshold vs chunk_count import matplotlib.pyplot as plt plt.plot(thresholds, chunk_counts) plt.xlabel("Similarity Threshold") plt.ylabel("Number of Chunks") plt.show()

Exercise 3: Compare Chunking Strategies

Goal: Implement all 5 strategies and compare on a real document.

Strategy Comparison

Implement fixed_size_chunking, recursive_chunking, semantic_chunking, document_aware_chunking, and agentic_chunking. For each, measure: number of chunks, average chunk size, time to chunk.

Python โ€” Starter Code
import time strategies = { "fixed": (fixed_size_chunking, {}), "recursive": (recursive_chunking, {}), "semantic": (semantic_chunking, {"embedding_func": get_embedding}), "doc_aware": (document_aware_chunking, {"doc_type": "html"}), } results = {} for name, (func, kwargs) in strategies.items(): start = time.time() chunks = func(document, **kwargs) elapsed = time.time() - start results[name] = { "num_chunks": len(chunks), "avg_size": sum(len(c.split()) for c in chunks) / len(chunks), "time_sec": elapsed } print(f"{name:12}: {len(chunks):3} chunks, " f"avg {results[name]['avg_size']:6.1f} words, " f"time {elapsed:.3f}s") # Which strategy is fastest? Most balanced?

Exercise 4: Chunk Quality Analysis

Goal: Develop metrics to measure chunk quality.

Quality Metrics

Compute for each chunk: coherence (do sentences relate to each other?), completeness (do sentences end properly?), length distribution. Identify low-quality chunks.

Python โ€” Starter Code
def analyze_chunk_quality(chunks): metrics = {} for i, chunk in enumerate(chunks): # Completeness: does it end with period/question mark? is_complete = chunk.strip().endswith(('.', '!', '?')) # Coherence: analyze sentence transitions sentences = chunk.split('.') # TODO: Compute coherence score # Length: is it reasonable? words = len(chunk.split()) is_reasonable = 50 < words < 1000 metrics[i] = { "complete": is_complete, "coherence": 0.8, # TODO: compute "reasonable": is_reasonable, "quality": 0.9 # TODO: combine scores } return metrics # Analyze quality = analyze_chunk_quality(chunks) low_quality = [i for i, m in quality.items() if m["quality"] < 0.7] print(f"Low quality chunks: {len(low_quality)} out of {len(chunks)}")

Exercise 5: Production RAG Pipeline

Goal: Build a complete RAG system with optimized chunking and measure end-to-end quality.

Production RAG Setup

Build complete pipeline: load PDFs โ†’ chunk with your best strategy โ†’ embed โ†’ create vector index โ†’ evaluate on test queries using NDCG@5 and MRR metrics.

Python โ€” Starter Code
from llama_index.core import VectorStoreIndex, Document # Load documents documents = load_pdf_directory("./documents/") # Chunk with optimal strategy parser = SemanticSplitterNodeParser(...) nodes = parser.get_nodes_from_documents(documents) # Create index index = VectorStoreIndex(nodes) # Evaluate test_queries = [ {"text": "What is X?", "relevant_doc_ids": [1, 3, 5]}, {"text": "How does Y work?", "relevant_doc_ids": [2, 4]}, ] ndcg_scores = [] for query in test_queries: results = index.retrieve(query["text"], top_k=5) # TODO: Compute NDCG ndcg_scores.append(ndcg) print(f"Average NDCG@5: {sum(ndcg_scores)/len(ndcg_scores):.3f}")

Interview Questions on Chunking

Q1: Explain the tradeoff between chunk size and retrieval quality. ▼
Small chunks (128 tokens) have higher precision โ€” fewer irrelevant words per chunk. But they may lose context and require more embedding computations. Large chunks (1024+ tokens) preserve context but include noise. The sweet spot is usually 256-512 tokens, but this depends on your embedding model and query patterns. The only way to know is to test on your data.
Q2: Why is overlap important in chunking? ▼
Overlap ensures that information split across chunk boundaries appears in multiple chunks. Without overlap, a question aligned with a chunk boundary might miss relevant information. With 20% overlap, important facts are represented in both the boundary chunk and the previous chunk, improving retrieval probability. The cost is ~2x more total chunks, but retrieval quality improvement justifies it.
Q3: Compare fixed-size vs. recursive chunking. When would you use each? ▼
Fixed-size is simplest and fastest, good for prototyping. Recursive respects document structure (paragraphs > sentences > words), better for heterogeneous documents. Use fixed-size for homogeneous plain text or rapid prototyping. Use recursive for documentation, code, markdown โ€” anything with natural hierarchy. Recursive adds negligible latency overhead.
Q4: What are the advantages and costs of semantic chunking? ▼
Semantic chunking uses embeddings to find natural topic boundaries, yielding 85-95% retrieval accuracy vs 75-85% for recursive. Cost: must embed every sentence (expensive at scale, adds latency). Best for mixed documents where structure isn't reliable. Worth it if retrieval quality is critical. Not worth it for simple documents with clear structure.
Q5: How would you handle chunking for different document types (PDF, HTML, code)? ▼
Different document types have different structure: PDF has pages, HTML has DOM hierarchy, code has classes/functions. Use format-specific parsers (PyPDF2 for PDF, BeautifulSoup for HTML) to extract text while preserving structure. Then apply chunking that respects that structure. LangChain and LlamaIndex provide document loaders that handle this. Always preserve metadata (page, section, line number) in chunks.
Q6: What is agentic chunking and when would you use it? ▼
Agentic chunking uses an LLM to read documents and decide how to chunk them based on semantic meaning. Highest quality (95%+ retrieval accuracy) but expensive ($0.30-1.00 per 100K tokens) and slow (seconds to minutes). Use only when document quality is critical (legal, medical, financial) and cost/latency are acceptable. Not for high-volume processing with tight budgets.
Q7: You're building a RAG system and notice retrieval quality is 65%. How would you debug and improve it? ▼
First: measure chunking quality independently. Retrieve chunks for a test query and manually score relevance. If chunks are relevant but LLM answer is poor, problem is LLM, not chunking. If chunks are irrelevant: (1) test different chunk sizes via NDCG experiments, (2) switch to semantic chunking if using fixed-size, (3) ensure metadata is preserved so filtering works. Run A/B tests with different strategies and pick the winner.
Q8: Explain how you would optimize chunk size for a specific use case. ▼
Run experiments: test chunk sizes 128, 256, 512, 1024, 2048. For each: create chunks, embed all chunks, evaluate on 50-100 representative test queries. Measure NDCG@5, MRR, Recall@5. Also measure latency and cost (num chunks * embedding cost). Plot NDCG vs chunk size and pick the size with best NDCG, or best quality/cost tradeoff. Typical sweet spot: 256-512 tokens.

Frequently Asked Questions

What's the difference between chunking and tokenization? ▼
Tokenization breaks text into tokens (subword units), usually for LLMs. A token might be a word like 'transformer' or a subword like 're', 'form', 'ing'. Chunking creates semantically coherent units (sentences, paragraphs, sections) that may contain hundreds of tokens. Chunking happens before indexing; tokenization happens inside the embedding/LLM model.
Does chunk size affect embedding quality? ▼
Yes. Embedding models (sentence transformers, OpenAI embeddings) are optimized for certain input lengths. text-embedding-3-small works best with 256-512 tokens. If you send much shorter text, it may pad and be less discriminative. If you send much longer text (>2000 tokens), it may lose semantic coherence. Always test your embedding model's optimal input length.
How much should chunks overlap? ▼
20-25% overlap is standard. This means a 512-token chunk overlaps 102-128 tokens with the previous chunk. This ensures context continuity at boundaries. Too little overlap (<10%) loses context. Too much overlap (>50%) creates excessive redundancy without quality gains. The right amount depends on your data โ€” test and measure.
Should I chunk before or after extracting metadata? ▼
Extract and attach metadata first, then chunk. This ensures each chunk knows its source document, page, section, etc. If you chunk first, you lose track of structure. Process: parse document โ†’ extract metadata (title, sections, page numbers) โ†’ chunk while preserving metadata โ†’ embed chunks โ†’ index.
How do I handle documents that are 100K tokens long? ▼
Don't try to process in one pass. Split into smaller subdocuments first (e.g., by chapter or section). Chunk each subdocument independently. Preserve the hierarchy in metadata. Alternatively, use hierarchical chunking: create summary chunks at document level, then detailed chunks at paragraph level. At query time, use multi-stage retrieval: search summaries first, then detailed chunks.
Can I chunk documents in different languages differently? ▼
Yes, and you should. Languages have different average word lengths (Chinese characters vs English words), different punctuation conventions, different grammar. For multilingual documents: detect language per document, use language-specific tokenizers, potentially use different chunk sizes per language. LangChain and LlamaIndex support this with language-aware splitters.
How often should I re-chunk documents? ▼
When documents are updated, re-chunk them. When you change chunking strategy, re-chunk everything (versioning important). If using semantic chunking with embeddings, re-chunking also means re-embedding (cost!). Common approach: chunk documents once when ingested, version the chunks. When strategy changes, create v2, v3, etc. At query time, you can A/B test different versions.
What's the relationship between chunk size and context window for the LLM? ▼
If your LLM has a 4K token context window and you retrieve 5 chunks of 512 tokens, that's 2560 tokens of context plus prompt overhead. Chunks leave room for prompt instructions and the user question. Don't make chunks so large that retrieved chunks consume most of the context window. Leave at least 25% of context window for LLM to reason.
Should I include chunk overlap in the embedding, or just in the index? ▼
Include overlap in the index. When you embed chunks, embed the full chunk including the overlap. This is correct because overlapping sentences should be represented in both embeddings. The overlap isn't a separate embedding โ€” it's part of both chunks' embeddings.
How do I test if my chunking is good? ▼
Measure retrieval quality: run test queries, retrieve top-5 chunks, score relevance (1 = relevant, 0 = irrelevant), compute NDCG@5. Target >0.85 NDCG. Also track: MRR (how fast you find first relevant), Recall@5 (what fraction of relevant chunks appear), latency, cost. Collect human judgments on a small sample to calibrate automated metrics.

Summary and Next Steps

Key Takeaways

Chunking is Foundational

It's the first step in RAG. Small chunking mistakes compound through embedding, retrieval, and generation. Get it right.

5 Main Strategies

Fixed-size (simple), Recursive (structured), Semantic (high-quality), Document-Aware (format-specific), Agentic (best quality). Choose based on your use case.

Test Your Assumptions

The 'best' chunk size isn't universal. Test 128, 256, 512, 1024 on your data. Measure NDCG@5 and MRR. Use winners.

Preserve Everything

Keep metadata: source document, page, section, character offsets. This enables proper attribution and filtering.

Decision Framework: Which Strategy?

Fast Prototyping?

Use Fixed-Size. Zero overhead, instant results. Good enough for MVP.

Structured Documents?

Use Recursive. Respects hierarchy. Fast and effective.

High Quality Needed?

Use Semantic. Embedding cost worth retrieval gains.

Critical Documents?

Use Agentic. LLM understands complex semantics.

Next Steps

  1. Start with Recursive: Good default for most use cases. Structure-aware, fast, free.
  2. Benchmark Your Data: Run optimization experiments. Test chunk sizes 128-2048. Pick winner by NDCG.
  3. Implement Evaluation: Create test query set with human-judged relevance. Measure NDCG@5, MRR continuously.
  4. Consider Semantic: If recursive NDCG < 0.80, upgrade to semantic chunking with threshold tuning.
  5. Production Ready: Implement versioning, monitoring, re-chunking for updates. Track retrieval quality in production.
  6. Optimize Iteratively: As you gather production data, re-run optimization. Chunk size needs may change.

Further Resources

Final Insight: Chunking isn't a solved problem. It's domain-specific, data-specific, and model-specific. The teams building the best RAG systems invest heavily in chunking optimization because it's one of the highest-ROI improvements available. Even small NDCG gains (0.80 โ†’ 0.85) significantly improve answer quality and user satisfaction.