Sections

Vector Databases: Comprehensive Guide

Master vector storage, similarity search, HNSW, IVF, and production platforms.

Vector Databases: Mastering Similarity Search

Vector databases have become essential infrastructure for modern AI applications. They enable efficient similarity search across billions of vectors, powering semantic search, recommendation systems, RAG (Retrieval-Augmented Generation), and AI-powered personalization at scale. While traditional databases excel at exact matching (SQL: WHERE id = 42), vector databases excel at finding similar items (Vector: FIND closest vectors to [0.1, 0.5, 0.3, ...]).

A vector database stores high-dimensional numerical representations (embeddings) and retrieves the most similar vectors using specialized indexing structures. Instead of comparing strings character-by-character or numbers with exact equality, vector databases compute similarity using distance metrics like cosine similarity or Euclidean distance.

What You'll Learn

Vector Fundamentals

Understand embeddings, dimensionality, and why vectors enable semantic understanding. Learn how language models create meaningful numerical representations of concepts.

Similarity Metrics

Master distance metrics: cosine similarity, Euclidean distance, Manhattan distance. Learn when to use each and how they impact performance.

Index Structures

Explore HNSW (Hierarchical Navigable Small World), IVF (Inverted File Index), LSH, and more. Understand time-space tradeoffs.

Production Systems

Deep dive into Pinecone, Qdrant, Weaviate, ChromaDB, Milvus, and pgvector. Learn when to use which platform.

Prerequisites

Basic Python, understanding of machine learning concepts, familiarity with embeddings/vectors. No advanced mathematics required โ€” we'll explain concepts from first principles.

Why Vector Databases Matter

The explosion of large language models (LLMs) and AI applications created new demands on data infrastructure. Traditional databases couldn't answer questions like "find the 10 most similar documents" efficiently. Vector databases solve this fundamental problem.

The Scale of the Opportunity

LLM Inference Latency (ms)
45ms
Vector DB Query (100M vectors)
125ms
SQL DB Full Scan (10M rows)
5234ms
Brute Force Cosine (1M vectors)
8932ms

Query latency for finding top-10 similar items โ€” lower is better

Key Use Cases

Semantic Search

Users search by meaning, not keywords. 'battery-powered transportation' returns bicycles AND electric scooters. Vector search understands intent.

RAG Systems

Retrieve relevant documents, then feed them to an LLM for generation. Powers ChatGPT plugins and enterprise AI assistants.

Recommendations

Find similar products/content/users at scale. Netflix recommends movies; TikTok recommends videos โ€” both using vector similarity.

Anomaly Detection

Outliers are vectors far from the normal cluster. Use vector distance to identify fraud, system failures, and unusual patterns.

The Vector Database Moment: Before 2023, vector search was a niche technique. After ChatGPT exploded and LLMs became mainstream, every developer suddenly needed semantic search. Vector databases went from 'nice to have' to 'table stakes' for modern applications.

Vector Fundamentals

Before diving into databases, let's understand vectors themselves. A vector is just an ordered list of numbers, but the key insight is that vectors can represent meaning.

From Text to Numbers

Language models (BERT, GPT, etc.) convert text into vectors called embeddings. Each dimension captures some aspect of meaning:

Python โ€” Creating Embeddings
from sentence_transformers import SentenceTransformer model = SentenceTransformer('all-MiniLM-L6-v2') # Text โ†’ embedding (384-dim vector) text1 = "A fast car" vector1 = model.encode(text1) # [0.123, -0.456, 0.789, ...] text2 = "A quick automobile" vector2 = model.encode(text2) # Similar vector to vector1 # High-level concept: similar meanings โ†’ similar vectors print(f"Vector dimension: {len(vector1)}") # Output: 384 print(f"Vector values: {vector1[:5]}") # [-0.12, 0.45, ...]

The beauty is that similar text produces similar vectors. Distance between vectors represents semantic difference.

Vector Space Properties

Vector Magnitude (L2 Norm)
||v|| = โˆš(vโ‚ยฒ + vโ‚‚ยฒ + ... + vโ‚™ยฒ) Measures the 'length' or 'magnitude' of a vector in n-dimensional space. Used in normalization: v_normalized = v / ||v||
Vector Dimensionality
Most embeddings are 300-3,000 dimensions: โ€ข BERT: 768 dimensions โ€ข GPT-3: 12,288 dimensions โ€ข OpenAI text-embedding-3-small: 512 dimensions โ€ข MiniLM: 384 dimensions Higher dimensions = more expressive but slower similarity search

Dimensionality Paradox

More dimensions provide richer representation, but also create the 'curse of dimensionality' โ€” distances become less meaningful, and search becomes slower. Often you must trade off representation quality vs. query speed.

Vector Normalization

Many systems normalize vectors (scale them to unit length). This matters because it affects distance calculations:

Python โ€” Vector Normalization
import numpy as np v = np.array([3, 4]) magnitude = np.linalg.norm(v) # โˆš(9 + 16) = 5.0 v_normalized = v / magnitude # [0.6, 0.8] # Check normalization print(np.linalg.norm(v_normalized)) # Output: 1.0 # Why normalize? For cosine similarity, normalized vectors # make the computation more efficient: # cos_sim(u, v) = u ยท v (only need dot product if normalized)

Vector Sparsity

Text embeddings are typically dense (all dimensions have non-zero values). Some systems use sparse vectors (only a few non-zero dimensions), which are faster but less expressive. SPLADE uses sparse vectors for fast, interpretable retrieval.

Similarity Search Fundamentals

Given a query vector, find the K most similar vectors from a database. This is a fundamental operation in all vector systems.

Exact vs. Approximate Search

Exact Search (Brute Force)

Compare query to every vector, compute exact distances, return top-K. Guaranteed correct answers. Time complexity: O(nยทd) where n = vectors, d = dimensions. For millions of vectors, this is too slow.

Approximate Search (Indexed)

Use index structure (HNSW, IVF, LSH) to skip comparisons. Trade small accuracy loss for massive speed gain. Usually find 99%+ of true neighbors in 1/100th the time.

Search Quality Metrics

Recall@K
Of the K results we returned, what percentage were actual top-K neighbors? Recall@10 = (# of true top-10 neighbors we found) / 10 Example: If we return 10 results and 9 are in the true top-10, Recall@10 = 9/10 = 90% Target: 99%+ recall with sub-10ms latency
Mean Average Precision (MAP)
Average precision across multiple queries: MAP = (1/Q) ฮฃ AP(q) for each query q AP = (1/K) ฮฃ (precision@i ร— is_relevant(i)) Penalizes mistakes in top positions more than bottom

The key insight: most applications accept 99% recall to get 100x speedup. A product recommendation that returns 99 of the 100 best matches is still excellent.

Recall vs Latency Tradeoff

Vector indices let you control this tradeoff. Index tuning parameters (ef_construction, M in HNSW, or n_probe in IVF) control accuracy/speed. More parameters = slower indexing but better recall, or faster queries but lower recall.

Similarity Search Example

Python โ€” Brute Force Search
import numpy as np # Database: 1000 vectors, 384 dimensions db_vectors = np.random.randn(1000, 384).astype('float32') # Normalize for cosine similarity db_vectors = db_vectors / np.linalg.norm(db_vectors, axis=1, keepdims=True) # Query vector query = np.random.randn(384).astype('float32') query = query / np.linalg.norm(query) # Brute force: compute distance to all vectors # Cosine similarity = dot product (after normalization) distances = db_vectors @ query # O(n*d) # Get top 10 most similar top_k = 10 top_indices = np.argsort(-distances)[:top_k] top_scores = distances[top_indices] print(f"Top {top_k} most similar:") for idx, score in zip(top_indices, top_scores): print(f" Vector {idx}: similarity {score:.4f}")

Distance Metrics for Vectors

Different metrics measure "distance" differently. Choosing the right metric is crucial for both correctness and performance.

Cosine Similarity

Cosine Similarity
cos_sim(u, v) = (u ยท v) / (||u|| ร— ||v||) Range: [-1, 1] where: 1 = identical vectors 0 = orthogonal (unrelated) -1 = opposite vectors Advantages: Fast, normalized, works in high dimensions Disadvantages: Ignores magnitude, sensitive to sparse vectors
Python โ€” Cosine Similarity
from sklearn.metrics.pairwise import cosine_similarity import numpy as np u = np.array([[1, 2, 3]]) v = np.array([[4, 5, 6]]) similarity = cosine_similarity(u, v)[0][0] print(f"Cosine similarity: {similarity:.4f}") # 0.9746 # For normalized vectors, cosine sim = dot product u_norm = u / np.linalg.norm(u) v_norm = v / np.linalg.norm(v) dot_product = (u_norm @ v_norm.T)[0][0] print(f"Dot product (normalized): {dot_product:.4f}") # 0.9746

Euclidean Distance

Euclidean Distance (L2)
d(u, v) = โˆš((uโ‚-vโ‚)ยฒ + (uโ‚‚-vโ‚‚)ยฒ + ... + (uโ‚™-vโ‚™)ยฒ) Range: [0, โˆž) where: 0 = identical vectors larger = more different Advantages: Intuitive, captures magnitude differences Disadvantages: Slower in high dimensions, magnitude-dependent
Python โ€” Euclidean Distance
from scipy.spatial.distance import euclidean import numpy as np u = np.array([1, 2, 3]) v = np.array([4, 5, 6]) distance = euclidean(u, v) print(f"Euclidean distance: {distance:.4f}") # 5.1962 # Also: L2 = sqrt(3), L2^2 = 3 l2_squared = np.sum((u - v) ** 2) print(f"L2 squared: {l2_squared}") # 27

Manhattan Distance

Manhattan Distance (L1)
d(u, v) = |uโ‚-vโ‚| + |uโ‚‚-vโ‚‚| + ... + |uโ‚™-vโ‚™| Range: [0, โˆž) Advantages: Fast (sum of absolute values), sparse-friendly Disadvantages: Less intuitive, doesn't account for direction

Hamming Distance

Hamming Distance
d(u, v) = count(uแตข โ‰  vแตข) for i = 1 to n Used for binary/discrete vectors (often quantized for speed) Example: u = [1, 0, 1, 1] v = [1, 1, 0, 1] Hamming distance = 2 (differ at positions 1 and 2)

Which Metric to Use?

Metric Best For Speed Space
Cosine Similarity Text/embeddings (ignore magnitude) Very Fast Requires normalization
Euclidean (L2) Geometric data, images Fast Accounts for all differences
Manhattan (L1) Sparse data, robustness Very Fast Lightweight
Hamming Binary/quantized vectors Fastest 1 bit per dimension

Metric Selection Rule of Thumb

Use cosine similarity for text/NLP embeddings (it's what language models are trained with). Use Euclidean for general purpose similarity or when magnitude matters. Use Manhattan for robustness to outliers. Use Hamming only with quantized/binary vectors.

HNSW: Hierarchical Navigable Small World

HNSW is the most popular index structure for vector databases. It powers Pinecone, Weaviate, Qdrant, and Milvus. Understanding it is essential for production systems.

Core Idea: Small World Networks

HNSW is inspired by the "small world" phenomenon โ€” in a small world network, you can reach any node via a few hops by following local connections. Think of social networks: you can reach someone across the world through 6 degrees of separation.

Small World Property: Even in huge networks, the average path length is logarithmic. HNSW exploits this: instead of comparing a query to all N vectors, you navigate a graph following distances, reaching the nearest neighbor in O(log N) comparisons on average.

The Algorithm

HNSW Parameters
M = max number of connections per node (default 5-16) ef_construction = size of dynamic candidate list during insert (default 200) ef = size of dynamic candidate list during search (default 10-100) Rules: โ€ข Higher M โ†’ better recall but slower/more memory โ€ข Higher ef_construction โ†’ better index quality but slower build โ€ข Higher ef โ†’ better recall but slower queries
Python โ€” HNSW Insertion Algorithm
# Simplified HNSW insertion pseudocode def insert(new_vector, M=5, ef_construction=200): # 1. Determine layer assignment (random, exponential) L = random_level() # 2. Find nearest neighbors at each layer from top to bottom for layer in range(L, 0, -1): # Greedy search: find closest node in this layer nearest = greedy_search(new_vector, layer) # 3. At layer 0, insert with M nearest neighbors neighbors_layer_0 = search_neighbors(new_vector, M, ef_construction) add_bidirectional_links(new_vector, neighbors_layer_0) # 4. Prune neighbor connections if needed for node in neighbors_layer_0: if node.connections > M: prune_worst_connections(node, M)

Layered Architecture

HNSW creates multiple layers. Higher layers (sparser) enable fast long-distance navigation. Lower layers (denser) enable fine-grained local search:

Layer 0 (Detailed)

Dense graph with most connections. Comparisons happen here. Thousands of links enable precise nearest neighbor finding.

Layers 1+ (Navigation)

Increasingly sparse. Few long-distance connections jump quickly toward the target region. Reduces the number of comparisons in lower layers.

Search Algorithm

Python โ€” HNSW Search
def search(query_vector, K=10, ef=10): nearest = [entry_point] # Start from top layer entry # Traverse layers top-to-bottom for layer in range(top_layer, 0, -1): # Greedy: move toward closest neighbor in this layer nearest = greedy_search(query_vector, nearest, layer) # At layer 0: beam search with ef candidates candidates = nearest for i in range(ef): # Explore neighbors of current candidates for neighbor in candidates.neighbors(layer=0): distance = compute_distance(query_vector, neighbor) candidates.add(neighbor, distance) # Return K nearest from final candidates return candidates.top_k(K)

HNSW Strengths and Weaknesses

Strengths: HNSW

Logarithmic query complexity, excellent recall with small ef values, in-memory friendly, handles insertions/deletions well, no reindexing needed.

Weaknesses: HNSW

Difficult to tune (M, ef_construction, ef), limited batch processing, random layer assignment adds variance, not ideal for disk-based systems.

HNSW in Practice

Start with M=16, ef_construction=200, ef=10. Measure recall@10 on test set. If recall < 95%, increase ef_construction or M. If latency > target, decrease ef or M. Typical P99 latency: 5-50ms for billions of vectors.

IVF: Inverted File Index

IVF (used in Faiss, Milvus, Vespa) partitions the vector space into clusters, then searches only the closest clusters. It trades recall for speed with larger indices.

Core Concept: Partitioning

IVF divides vectors into N clusters (typically 256-16,384). Each cluster has a centroid. To search:

  1. Find the n_probe closest clusters to the query
  2. Search only vectors in those clusters
  3. Return the top-K overall
IVF Time Complexity
Without index: O(n ร— d) With IVF: O(k + (n/c) ร— n_probe ร— d) Where: n = total vectors d = dimensions c = number of clusters k = top-K results n_probe = clusters to search (1-c) Typical speedup: 10-100x with small accuracy loss
Python โ€” IVF Search Example
# Simplified IVF search pseudocode def ivf_search(query_vector, K=10, n_probe=10): # 1. Find distance to all cluster centroids O(c*d) distances_to_centroids = [ compute_distance(query_vector, centroid) for centroid in centroids ] # 2. Find n_probe closest clusters closest_clusters = np.argsort(distances_to_centroids)[:n_probe] # 3. Search only vectors in those clusters candidates = [] for cluster_id in closest_clusters: cluster_vectors = get_cluster_vectors(cluster_id) for vector in cluster_vectors: distance = compute_distance(query_vector, vector) candidates.append((distance, vector)) # 4. Return top-K candidates.sort() return candidates[:K]

Building the Index

K-Means Clustering

IVF uses K-means to partition vectors. Run K-means on sample (10-100M vectors) to find centroids. Assign all vectors to nearest centroid. Build inverted list: for each centroid, store all vectors in that cluster.

Tuning IVF: The Recall/Speed Tradeoff

n_probe=1
5ms
n_probe=10
35ms
n_probe=50
145ms
n_probe=256 (all)
1850ms

IVF latency vs recall tradeoff. Search 10M vectors, K=10

IVF vs HNSW

Aspect HNSW IVF
Index Size Smaller (fewer links) Larger (stores cluster assignments)
Building Time O(n log n), slower O(n log k), faster
Query Latency 5-20ms typical 5-50ms (variable)
Tuning Complexity Moderate (M, ef, ef_construction) High (n_clusters, n_probe, retraining)
Insertions/Updates Fast, no reindex Slow, may need reindex
Best For Streaming, dynamic data Batch operations, static data

ChromaDB: Simple Vector Database

ChromaDB is the easiest vector database to get started with. It's open-source, requires zero infrastructure, and works great for prototyping and small-to-medium scale applications.

Key Features

Easy Setup

pip install chromadb, import, use. No Docker, no databases, no credentials. Perfect for rapid prototyping.

Built-in Embeddings

Chromadb can generate embeddings automatically using sentence-transformers or OpenAI. Just add text, it handles the rest.

Multiple Backends

Run in-memory, on-disk, or connect to a Chroma server. Scale from local laptop to distributed setup.

Metadata Filtering

Filter results by metadata before/after similarity search. E.g., only documents from 2024.

Installation and Basic Usage

Python โ€” ChromaDB Basic Operations
# Installation # pip install chromadb import chromadb # Create a local persistent database client = chromadb.PersistentClient(path="/tmp/chroma") # Create a collection (like a table) collection = client.get_or_create_collection(name="documents") # Add documents (Chroma auto-generates embeddings) collection.add( ids=["doc1", "doc2", "doc3"], documents=[ "The quick brown fox jumps over the lazy dog", "A fast red fox leaps over a sleepy dog", "Python is a popular programming language" ], metadatas=[ {"source": "story.txt", "year": 2024}, {"source": "story.txt", "year": 2024}, {"source": "programming.txt", "year": 2023} ] ) # Query: find 2 most similar documents results = collection.query( query_texts=["A quick fox"], n_results=2 ) print(results["documents"]) # [["The quick brown...", "A fast red..."]] print(results["distances"]) # [[0.15, 0.32]] # Lower = more similar print(results["metadatas"]) # Metadata for results

Advanced: Filtering and Custom Embeddings

Python โ€” ChromaDB with Filtering
import chromadb from sentence_transformers import SentenceTransformer client = chromadb.PersistentClient(path="/tmp/chroma") collection = client.get_or_create_collection(name="articles") # Add documents with rich metadata collection.add( ids=["a1", "a2", "a3", "a4"], documents=[ "Machine learning is transforming AI", "Deep learning requires lots of data", "Python is easy to learn", "Rust is fast and safe" ], metadatas=[ {"category": "AI", "year": 2024}, {"category": "AI", "year": 2023}, {"category": "Programming", "year": 2024}, {"category": "Programming", "year": 2024} ] ) # Query with metadata filter: only AI articles from 2024 results = collection.query( query_texts=["neural networks"], n_results=2, where={"$and": [ {"category": "AI"}, {"year": 2024} ]} ) # Custom embeddings with sentence-transformers model = SentenceTransformer('all-mpnet-base-v2') embeddings = [model.encode(doc).tolist() for doc in my_documents] collection2 = client.get_or_create_collection(name="custom") collection2.add( ids=list(range(len(my_documents))), documents=my_documents, embeddings=embeddings )

When to Use ChromaDB

Perfect For:

Prototyping RAG systems, small document stores (< 100K), learning vector databases, demo applications, hobby projects.

Not For:

Production systems with >1M documents, heavy concurrent access, sub-5ms latency requirements, complex filtering, multi-tenant.

ChromaDB Performance

In-memory: 1-5ms per query (up to 100K vectors). Persistent: 5-20ms. Server mode: 20-100ms (includes network). Suitable for applications where 50-100ms latency is acceptable.

Pinecone: Managed Vector Database

Pinecone is the most mature managed vector database service. It handles infrastructure, scaling, and availability so you just push vectors and query. Perfect for production AI applications.

Key Features

Fully Managed

No ops, no infrastructure. Just use API. Automatic scaling, replication, disaster recovery built-in.

Production Ready

99.9% SLA, low-latency endpoints, edge caching, real-time indexing. Handles millions of vectors easily.

Pod Storage

Persistent storage for metadata. Store arbitrary JSON with each vector for rich filtering and retrieval.

Serverless

Pay-per-request pricing. No minimum costs. Scales from zero to billions of vectors.

Setup and Basic Operations

Python โ€” Pinecone Setup and Search
# pip install pinecone-client from pinecone import Pinecone import openai # Initialize pc = Pinecone(api_key="YOUR_API_KEY") index = pc.Index("my-index") # Upload vectors vectors_to_upsert = [ ("vec1", [0.1, 0.2, 0.3, ...], {"text": "First document", "date": "2024-01-01"}), ("vec2", [0.2, 0.3, 0.4, ...], {"text": "Second document", "date": "2024-01-02"}), ] index.upsert(vectors=vectors_to_upsert) # Query query_embedding = openai.Embedding.create( input="machine learning tutorial", model="text-embedding-3-small" )["data"][0]["embedding"] results = index.query( vector=query_embedding, top_k=10, include_metadata=True ) for match in results.matches: print(f"{match.metadata['text']}: {match.score}")

Advanced: Hybrid Search with Metadata

Python โ€” Pinecone with Filtering
from pinecone import Pinecone pc = Pinecone(api_key="YOUR_API_KEY") index = pc.Index("articles") # Upsert with metadata index.upsert(vectors=[ ("id1", vector1, { "title": "AI Trends 2024", "category": "AI", "views": 5000, "year": 2024 }), ("id2", vector2, { "title": "Python Async", "category": "Programming", "views": 2000, "year": 2024 }), ]) # Query with metadata filter (namespace support) results = index.query( vector=query_vec, top_k=5, filter={ "$and": [ {"category": {"$eq": "AI"}}, {"year": {"$gte": 2024}} ] }, include_metadata=True ) # Hybrid search: combine vector similarity + filters # Returns vectors similar to query that also match filters

Performance Characteristics

P50 Latency
8ms
P95 Latency
25ms
P99 Latency
60ms
Throughput (QPS)
10000+

Typical performance on standard index (billions of vectors)

Pricing Model

Pinecone Costs

Serverless: $0.04 per 1M requests. Standard indexes: $0.25/pod-hour (s1 pod) to $2.50/pod-hour (p2 pod). Each pod = ~2M vectors (s1) to ~40M vectors (p2). For 10M vectors with 100 QPS: ~$200-500/month.

When to Use Pinecone

Best For:

Production AI apps, RAG systems at scale, SaaS products, teams without ops expertise, applications needing 99.9% uptime.

Trade-offs:

Vendor lock-in, pricing can add up, limited customization of index internals, requires API key management.

Qdrant: Open Source, Powerful Filtering

Qdrant is an open-source vector database excelling at rich metadata filtering. It combines the best of HNSW indexing with SQL-like filtering capabilities, making it ideal for complex retrieval scenarios.

Key Features

Open Source

Full control, no vendor lock-in, can self-host or use managed cloud, MIT licensed.

Advanced Filtering

Filter on metadata using complex queries (AND, OR, nested conditions). Unlike simpler systems, Qdrant filters efficiently.

Payload Storage

Store rich JSON metadata with vectors. Query using complex conditions: {age > 18 AND category IN ['ai', 'ml']}

Python/Rust

High-performance Rust backend. Python-first API. Cloud version available with SLA.

Installation and Basic Operations

Python โ€” Qdrant Setup
# pip install qdrant-client from qdrant_client import QdrantClient from qdrant_client.models import Distance, VectorParams, PointStruct # Connect to Qdrant (local or remote) client = QdrantClient(":memory:") # In-memory for testing # Or: QdrantClient("localhost", port=6333) # Docker instance # Or: QdrantClient(url="https://...", api_key="...") # Cloud # Create collection with vector config client.create_collection( collection_name="documents", vectors_config=VectorParams(size=384, distance=Distance.COSINE) ) # Upsert points (vectors with metadata) points = [ PointStruct( id=1, vector=[0.1, 0.2, 0.3, ...], # 384-dimensional vector payload={ "text": "Machine learning fundamentals", "category": "AI", "date": "2024-01-15", "rating": 4.5 } ), PointStruct( id=2, vector=[0.15, 0.25, 0.35, ...], payload={ "text": "Python for data science", "category": "Programming", "date": "2024-02-01", "rating": 4.8 } ), ] client.upsert( collection_name="documents", points=points ) # Simple search results = client.search( collection_name="documents", query_vector=[0.12, 0.22, 0.32, ...], # Query vector limit=5 ) for result in results: print(f"ID: {result.id}, Score: {result.score}, Text: {result.payload['text']}")

Advanced: Complex Metadata Filtering

Python โ€” Qdrant with Metadata Filtering
from qdrant_client.models import PointStruct, Filter, FieldCondition, MatchValue, Range # Search with complex filter results = client.search( collection_name="documents", query_vector=query_vec, query_filter=Filter( must=[ FieldCondition( key="category", match=MatchValue(value="AI") ), FieldCondition( key="rating", range=Range(gte=4.0) ) ], must_not=[ FieldCondition( key="date", range=Range(lte=1672531200) # Before 2023 ) ] ), limit=10 ) # Complex query examples: # 1. (category='AI' OR category='ML') AND rating >= 4.5 AND date > 2024-01-01 # 2. has_images=true AND location within_radius(10km) # 3. price < 100 AND stock > 0 AND (brand='Apple' OR brand='Samsung')

Qdrant Performance on Filtering

Simple search (no filter)
2ms
Search + filter (10% selectivity)
3ms
Search + complex filter (5% selectivity)
8ms
Search + filter (0.1% selectivity)
45ms

Latency with increasing filter selectivity (10M vectors)

Docker Setup

Python โ€” Run Qdrant with Docker
# Start Qdrant server docker run -p 6333:6333 qdrant/qdrant # Connect from Python from qdrant_client import QdrantClient client = QdrantClient("localhost", port=6333) # Now use as normal

When to Use Qdrant

Best For:

Applications needing complex filtering, self-hosted deployments, open-source preference, hybrid retrieval (vector + structured search).

Trade-offs:

More complex to set up than Pinecone, less mature ecosystem, requires self-hosting or managing cloud deployment.

Weaviate: GraphQL-Based Vector Database

Weaviate combines vector search with graph capabilities and structured data. It's designed for complex knowledge graphs and enterprises needing both semantic and structured search.

Key Features

GraphQL API

Query using GraphQL instead of REST. Single query combines vector search, filtering, and relationships.

Hybrid Search

Combine vector similarity with keyword (BM25) search. Best of both worlds for comprehensive retrieval.

Multi-Tenancy

Built-in multi-tenant support. Perfect for SaaS platforms with isolated customer data.

Built-in Transformers

Vectorize text automatically using Hugging Face models. No external embedding service needed.

Setup and Basic Usage

Python โ€” Weaviate Setup with Python
# pip install weaviate-client import weaviate from weaviate.classes.generic import DataObject # Connect to Weaviate client = weaviate.connect_to_local() # Docker required # Define schema (class structure) articles_schema = { "classes": [ { "class": "Article", "description": "A news article", "properties": [ { "name": "title", "dataType": ["text"], "description": "Article title", }, { "name": "content", "dataType": ["text"], "description": "Article content", }, { "name": "category", "dataType": ["text"], "description": "Category (AI, Tech, etc)", }, { "name": "publication_date", "dataType": ["date"], "description": "When published", }, ], "vectorizer": "text2vec-openai", # Auto-vectorize with OpenAI "moduleConfig": { "text2vec-openai": { "model": "text-embedding-3-small" } } } ] } client.schema.create(articles_schema) # Add data (auto-vectorized) client.data.create( collection_name="Article", properties={ "title": "Deep Learning Revolution", "content": "Recent advances in deep learning...", "category": "AI", "publication_date": "2024-02-27" } ) # Query using GraphQL response = client.graphql.raw(""" { Get { Article( nearText: {concepts: ["machine learning trends"]} where: { operator: AND operands: [ {path: ["category"], operator: Equal, valueString: "AI"} {path: ["publication_date"], operator: GreaterOrEqual, valueDate: "2024-01-01"} ] } limit: 10 ) { title content category publication_date _additional { distance } } } } """) print(response)

Hybrid Search (Vector + Keyword)

Python โ€” Weaviate Hybrid Search
# Hybrid search combines vector + BM25 keyword search response = client.graphql.raw(""" { Get { Article( hybrid: { query: "transformer attention mechanism", alpha: 0.75 # 75% vector, 25% keyword } limit: 5 ) { title _additional { score explainScore } } } } """) # alpha = 0.0: Pure keyword (BM25) # alpha = 0.5: Balanced vector + keyword # alpha = 1.0: Pure vector similarity # Typical: alpha = 0.75 (mostly semantic, some keyword)

Docker Setup

Python โ€” Run Weaviate
# Start with docker-compose version: '3.4' services: weaviate: image: semitechnologies/weaviate:latest ports: - "8080:8080" environment: QUERY_DEFAULTS_LIMIT: 25 AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: "true" PERSISTENCE_DATA_PATH: "/var/lib/weaviate" # Start: docker-compose up -d

When to Use Weaviate

Best For:

Knowledge graphs, enterprises, hybrid search needs, multi-tenant SaaS, semantic web applications.

Trade-offs:

Steeper learning curve (GraphQL), heavier resource usage, slower to set up than simpler alternatives.

pgvector: PostgreSQL Vector Extension

pgvector adds vector search to PostgreSQL. Perfect if you already use Postgres and want to avoid a separate database. Great for integrating vector search into existing applications.

Key Features

PostgreSQL Native

Vector type directly in Postgres. Use SQL familiar to millions of developers. No new infrastructure.

Multiple Indexes

IVFFLAT (IVF) and HNSW indexes. Choose based on your workload.

SQL Integration

Combine vector search with any SQL query. Join vectors with relational data seamlessly.

ACID Compliance

Transactions, consistency guarantees. Vectors are first-class citizens in Postgres.

Installation and Setup

Python โ€” pgvector Installation
# On the server (assuming Ubuntu) sudo apt-get install postgresql-contrib sudo apt-get install build-essential postgresql-server-dev-14 git clone --branch v0.5.1 https://github.com/pgvector/pgvector.git cd pgvector make sudo make install # In PostgreSQL: CREATE EXTENSION vector; # Verify SELECT * FROM pg_extension WHERE extname = 'vector';

Basic Vector Operations

Python โ€” pgvector with Python
import psycopg2 import numpy as np from pgvector.psycopg2 import register_vector # Connect and register vector type conn = psycopg2.connect("dbname=mydb user=postgres") register_vector(conn) cur = conn.cursor() # Create table with vector column cur.execute(""" CREATE TABLE IF NOT EXISTS documents ( id SERIAL PRIMARY KEY, title TEXT NOT NULL, content TEXT NOT NULL, embedding vector(384), -- 384-dim vector category TEXT, created_at TIMESTAMP DEFAULT CURRENT_TIMESTAMP ); """) # Create index (HNSW recommended) cur.execute(""" CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops) WITH (m=16, ef_construction=64); """) # Insert documents with embeddings embedding = np.random.randn(384).astype('float32') cur.execute( "INSERT INTO documents (title, content, embedding, category) VALUES (%s, %s, %s, %s)", ("ML Tutorial", "Learn machine learning...", embedding.tolist(), "AI") ) conn.commit() # Vector similarity search (cosine) cur.execute(""" SELECT id, title, 1 - (embedding <=> %s) as similarity FROM documents WHERE category = 'AI' ORDER BY embedding <=> %s LIMIT 10; """, (query_embedding, query_embedding)) results = cur.fetchall() for id, title, similarity in results: print(f"{id}: {title} (similarity: {similarity:.4f})") cur.close() conn.close()

Vector Operators

Operator Meaning Example
<-> Euclidean distance embedding <-> query_vec
<#> Negative dot product embedding <#> query_vec
<=> Cosine distance embedding <=> query_vec
<|> Dot product embedding <|> query_vec

Index Types: IVFFLAT vs HNSW

Python โ€” pgvector Index Options
-- IVFFLAT (Inverted File Flat) - simpler, less memory CREATE INDEX ON documents USING ivfflat (embedding vector_cosine_ops) WITH (lists=100); -- Number of clusters -- HNSW (Hierarchical Navigable Small World) - faster, more memory CREATE INDEX ON documents USING hnsw (embedding vector_cosine_ops) WITH (m=16, ef_construction=64); -- m = number of connections per node (16 is typical) -- ef_construction = size of search space during build (64-200) -- ef = size of search space during query (can set at query time) -- Query with specific ef parameter (HNSW only) SET hnsw.ef_search = 100; SELECT * FROM documents ORDER BY embedding <=> query_vec LIMIT 10;

Hybrid SQL + Vector Queries

Python โ€” Complex pgvector Queries
-- Find top 10 most similar documents -- published in last 30 days, in AI category SELECT id, title, 1 - (embedding <=> %s) as similarity, created_at FROM documents WHERE category = 'AI' AND created_at > NOW() - INTERVAL '30 days' ORDER BY embedding <=> %s LIMIT 10; -- Similarity search with aggregation SELECT category, AVG(1 - (embedding <=> %s)) as avg_similarity, COUNT(*) as count FROM documents GROUP BY category ORDER BY avg_similarity DESC; -- Find documents within distance threshold SELECT id, title FROM documents WHERE (embedding <=> %s) < 0.5; -- Distance < 0.5

Performance Tuning

pgvector Indexing Tips

Use HNSW for production (better latency). For HNSW, m=16 is usually optimal. ef_construction should be 200-400 for high-quality index. At query time, ef_search defaults to ef_construction but can be tuned (higher = more accurate but slower). For 1M vectors with HNSW: ~1-10ms queries.

When to Use pgvector

Best For:

Existing Postgres users, hybrid relational+vector queries, applications already using SQL, cost-conscious deployments.

Trade-offs:

Not specialized like Pinecone (slower), requires Postgres administration, limited built-in tools for embeddings.

Platform Comparison: Pinecone vs Qdrant vs Weaviate vs Milvus vs Chroma

Each platform has strengths and trade-offs. This comparison helps you choose the right one for your use case.

Feature Comparison Matrix

Feature Pinecone Qdrant Weaviate Milvus ChromaDB
Hosting Managed only Self-host or cloud Self-host or cloud Self-host only Local or server
Index Type HNSW HNSW HNSW/IVF HNSW/IVF HNSW
Metadata Filtering Basic Advanced Advanced Advanced Basic
API Type REST/gRPC REST/gRPC GraphQL/REST REST/Python/Go Python library
Max Vectors Billions Billions Billions Billions Millions
P99 Latency <100ms <100ms <200ms <500ms <50ms*
Learning Curve Easy Moderate Hard Moderate Very Easy
Setup Time <5 min 30 min 1 hour 1-2 hours <1 min
Cost High (managed) Low (self) / Medium (cloud) Medium Low Free

* ChromaDB: in-memory is fast but limited to millions. *Latencies depend on index size and configuration

Decision Matrix: Which to Choose?

Prototyping & Learning: ChromaDB. Zero setup, instant results, no costs. Perfect for learning vector databases.

Production SaaS/API: Pinecone or Weaviate. Managed infrastructure, high availability, focus on your code not ops.

Complex Filtering: Qdrant or Weaviate. Rich metadata queries, SQL-like filters, flexible data models.

Existing Postgres Users: pgvector. Minimal new infrastructure, leverage existing expertise and connections.

Cost-Conscious, Self-Hosting: Milvus or Qdrant. Open-source, control your costs, full customization.

Enterprise/Hybrid: Weaviate. Multi-tenancy, hybrid search (vector + keyword), knowledge graph capabilities.

Scaling Vector Databases

How do you scale from thousands to billions of vectors? Different platforms have different scaling models.

Scaling Dimensions

Data Size

From millions to billions of vectors. Sharding across nodes is necessary beyond single-machine capacity.

Throughput

From hundreds to millions of QPS (queries per second). Replication and caching needed for high QPS.

Latency

P50, P95, P99 tail latencies must remain low even under high load. Requires careful indexing and caching.

Updates

How frequently you add/delete/update vectors. Some systems excel at real-time updates, others prefer batch operations.

Sharding Strategy

Shard Assignment
Shard ID = hash(vector_id) % num_shards Each shard stores a subset of vectors. Queries are broadcast to all shards in parallel. Results are aggregated and top-K selected. Example: 1B vectors ร— 100 shards = 10M vectors per shard

Most managed systems (Pinecone, Weaviate Cloud) handle sharding automatically. Self-hosted systems (Qdrant, Milvus) require explicit configuration.

Replication for High Availability

Python โ€” Replication Example
# Typical setup: 3 replicas for fault tolerance # If one node fails, others handle traffic cluster_config = { "shard_count": 100, # 100 shards total "replication_factor": 3, # 3 copies of each shard "total_nodes": 10 # At least replication_factor nodes } # Distribution: # Total shards with replicas: 100 ร— 3 = 300 # Shards per node: 300 / 10 = 30 shards per node # Throughput scaling: # 1 shard @ 1000 QPS = 1000 QPS # 10 shards @ 1000 QPS each = 10,000 QPS # 100 shards = 100,000 QPS with replication

Caching Strategies

Multi-Level Caching

L1: In-memory node-local cache (hot vectors). L2: Distributed cache across replicas. L3: Disk-based persistent storage. Typically: 1% of data in L1, 10% in L2, rest on disk. This enables sub-10ms latency even for billion-scale systems.

L1 Cache Hit
<1ms
L2 Cache Hit
5-15ms
Disk Read
50-200ms

Partitioning Strategies

Strategy Approach Pros Cons
Hash-Based Shard = hash(ID) % N Simple, even distribution Rebalancing on growth is hard
Range-Based Shard = ID range Range queries efficient Can have hotspots
Space-Filling Curve Z-order (Morton) curve Locality-preserving, good cache behavior Complex to implement
Consistent Hashing Ring of hash values Minimal rebalancing on scale More overhead than hash

Optimization Techniques

Getting the best performance from vector databases requires careful tuning and optimization.

Quantization: Trading Accuracy for Speed

Quantization reduces vector dimensionality or bit-depth, trading small accuracy loss for large speed/memory gains.

Vector Quantization
Original: 384-dim float32 = 1536 bytes per vector 1. Product Quantization (PQ): Split into m segments, quantize each to k bits 384 dimensions โ†’ 8 segments ร— 4 bits = 4 bytes per vector 96x compression! Only ~0.1% recall loss 2. Binary Quantization: Each dimension โ†’ 1 bit (above/below median) 384-dim โ†’ 48 bytes โ†’ 32x smaller Find nearest in Hamming distance (very fast) 3. Scalar Quantization: float32 โ†’ int8: 4x smaller Often: use int8 for filtering, float32 for reranking
Python โ€” Quantization Example
import numpy as np # Original vectors (float32) vectors = np.random.randn(1000, 384).astype('float32') print(f"Original size: {vectors.nbytes / 1e6:.1f} MB") # 1.5 MB # Binary Quantization def binary_quantize(v): """Convert float32 to binary (1 bit per dimension)""" median = np.median(v) return (v > median).astype('uint8') binary_vecs = np.array([binary_quantize(v) for v in vectors]) print(f"Binary quantized: {binary_vecs.nbytes / 1e6:.3f} MB") # 0.048 MB print(f"Compression: {1.5 / 0.048:.0f}x smaller") # Product Quantization (simplified) def pq_quantize(v, num_subvectors=8): """Quantize to 8-bit per subvector""" subvector_size = len(v) // num_subvectors quantized = [] for i in range(num_subvectors): subvec = v[i*subvector_size:(i+1)*subvector_size] # Quantize subvector to 256 levels (8 bits) min_val, max_val = subvec.min(), subvec.max() normalized = (subvec - min_val) / (max_val - min_val + 1e-6) quantized.append(int(normalized.mean() * 255)) return np.array(quantized, dtype='uint8') pq_vecs = np.array([pq_quantize(v) for v in vectors]) print(f"PQ quantized: {pq_vecs.nbytes / 1e6:.2f} MB") # 0.008 MB print(f"Compression: {1.5 / 0.008:.0f}x smaller")

Caching and Prefetching

Smart Caching

Cache frequently accessed vector embeddings in memory. Use LRU (Least Recently Used) eviction. For datasets with skewed access (some documents are searched 100x more), in-memory caching can eliminate 70%+ of disk I/O.

Batch Search Optimization

When searching multiple queries simultaneously, batch them to amortize overhead and improve cache locality:

Python โ€” Batch vs Single Search
# Single query: 10ms overhead + 2ms compute results = [] for query in queries: results.append(index.search(query, k=10)) # Total: 100 queries ร— 12ms = 1200ms # Batch search: 10ms overhead once + 2ms per query (vectorized) batch_results = index.batch_search(queries, k=10) # Total: 10ms + 100ร—2ms = 210ms # ~6x faster! # How it works: # - Vectorized operations on GPU # - Better CPU cache behavior # - Amortized I/O operations # - Network round-trip batching (if remote)

Index Tuning Parameters

HNSW Tuning Guide
For Recall@10 >= 95% with minimum latency: 1. Index Building (one-time): ef_construction = 200-400 M = 12-16 2. Query Time: Start with ef = ef_construction/2 Measure recall, tune up if needed 3. Memory Tradeoff: Memory โ‰ˆ n_vectors ร— (d ร— 4 + M ร— 4 ร— 8) For 100M vectors, d=384, M=16: โ‰ˆ 100M ร— (1536 + 512) = 205 GB Consider quantization if memory-constrained

Benchmark Template

Python โ€” Benchmarking Vector Database
import time import numpy as np from statistics import mean, stdev def benchmark_search(db, queries, k=10, num_runs=3): """Benchmark search latency and recall""" latencies = [] recalls = [] for _ in range(num_runs): # Warm up db.search(queries[0], k=k) # Measure start = time.perf_counter() results = db.search(queries, k=k) latencies.append((time.perf_counter() - start) / len(queries)) avg_latency = mean(latencies) * 1000 # Convert to ms std_latency = stdev(latencies) * 1000 print(f"Latency: {avg_latency:.2f}ms ยฑ {std_latency:.2f}ms") print(f"Throughput: {1000/avg_latency:.0f} QPS") return avg_latency # Compare configurations for M in [8, 12, 16, 24]: for ef_construction in [200, 400, 600]: print(f"\nM={M}, ef_construction={ef_construction}") benchmark_search(db, test_queries)

Complete Code Examples

Building a Simple RAG System with Chroma

Python โ€” RAG with ChromaDB
import chromadb from sentence_transformers import SentenceTransformer import openai # 1. Setup client = chromadb.PersistentClient(path="./chroma_data") collection = client.get_or_create_collection("documents") embed_model = SentenceTransformer('all-MiniLM-L6-v2') # 2. Index documents documents = [ "Machine learning is a subset of AI focusing on learning from data.", "Deep learning uses neural networks with multiple layers.", "NLP processes and understands human language.", "Computer vision analyzes images and video." ] collection.add( ids=[str(i) for i in range(len(documents))], documents=documents, metadatas=[{"source": f"doc_{i}"} for i in range(len(documents))] ) # 3. RAG: Retrieve + Generate def rag_query(question): # Step 1: Retrieve relevant documents results = collection.query(query_texts=[question], n_results=2) context = "\n".join(results["documents"][0]) # Step 2: Generate answer using LLM response = openai.ChatCompletion.create( model="gpt-4", messages=[ { "role": "system", "content": "Answer based on the provided context." }, { "role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}" } ] ) return response["choices"][0]["message"]["content"] # Test answer = rag_query("What is deep learning?") print(answer)

Similarity Search with Pinecone

Python โ€” Pinecone Production Setup
from pinecone import Pinecone import openai import numpy as np # Initialize pc = Pinecone(api_key="your-api-key") index = pc.Index("documents") # Upsert millions of vectors efficiently def batch_upsert(documents, batch_size=100): """Efficiently upload large batches""" for i in range(0, len(documents), batch_size): batch = documents[i:i+batch_size] vectors = [ (doc['id'], openai.Embedding.create( input=doc['text'], model="text-embedding-3-small" )["data"][0]["embedding"], {"text": doc['text'], "date": doc['date']}) for doc in batch ] index.upsert(vectors=vectors) print(f"Upserted {i+len(batch)}/{len(documents)}") # Semantic search function def semantic_search(query_text, top_k=10, filters=None): # Embed query query_embedding = openai.Embedding.create( input=query_text, model="text-embedding-3-small" )["data"][0]["embedding"] # Search with optional filtering results = index.query( vector=query_embedding, top_k=top_k, filter=filters, include_metadata=True ) return results # Example usage results = semantic_search( "latest AI breakthroughs", top_k=5, filters={"date": {"$gte": "2024-01-01"}} ) for match in results.matches: print(f"{match.metadata['text']}: {match.score:.4f}")

Advanced Filtering with Qdrant

Python โ€” Qdrant Complex Filtering
from qdrant_client import QdrantClient from qdrant_client.models import Filter, FieldCondition, HasIdCondition, Range client = QdrantClient("localhost", port=6333) # Multi-level filtering example advanced_filter = Filter( must=[ # All of these must be true FieldCondition(key="product_category", match={"value": "electronics"}), FieldCondition(key="price", range=Range(gte=100, lte=1000)), ], should=[ # At least one of these should be true FieldCondition(key="brand", match={"value": "Apple"}), FieldCondition(key="brand", match={"value": "Samsung"}), ], must_not=[ # None of these should be true FieldCondition(key="out_of_stock", match={"value": True}), ] ) results = client.search( collection_name="products", query_vector=product_embedding, query_filter=advanced_filter, limit=20 ) for result in results: print(f"{result.payload['name']}: {result.score:.3f}")

Practical Exercises

Exercise 1: Build Your First Vector Database

Create a ChromaDB collection with 100 sample documents, then search for 5 queries. Measure latency.

Python โ€” Starter Code
import chromadb client = chromadb.EphemeralClient() collection = client.create_collection("docs") # TODO: Add 100 documents # TODO: Query 5 times # TODO: Measure and print latency

Exercise 2: Compare Distance Metrics

Generate 1000 random vectors. Measure query latency and recall for cosine vs Euclidean distance.

Python โ€” Starter Code
import numpy as np from sklearn.metrics.pairwise import cosine_similarity, euclidean_distances vectors = np.random.randn(1000, 384) query = np.random.randn(384) # TODO: Measure cosine vs euclidean # TODO: Which is faster? # TODO: Do results differ?

Exercise 3: Tune HNSW Parameters

Build an HNSW index with different M and ef_construction values. Plot recall vs latency tradeoff.

Python โ€” Starter Code
# Using Qdrant or similar import time configs = [ {"m": 8, "ef_construction": 200}, {"m": 16, "ef_construction": 400}, {"m": 32, "ef_construction": 800}, ] results = [] for config in configs: # Build index with config # Measure: latency, recall # Store in results # TODO: Plot recall vs latency # TODO: Which configuration is best?

Exercise 4: Build a RAG System

Create a system that retrieves documents based on semantic similarity, then answers questions using an LLM.

Python โ€” Starter Code
# Using ChromaDB + OpenAI documents = [...] # Your documents queries = [...] # Your questions # TODO: Index documents in ChromaDB # TODO: For each query: # - Retrieve top 3 most similar documents # - Use GPT-4 to generate answer from context # - Print answer # TODO: Evaluate: Is the answer correct?

Interview Questions & Answers

Q1: What's the difference between exact and approximate nearest neighbor search? ▼
Answer: Exact search compares the query to every vector and returns the true K nearest. Approximate search uses an index (like HNSW) to skip comparisons, trading small accuracy loss (typically 1-5% recall loss) for massive speed gains (10-100x faster). In practice, most applications accept 99% recall for 100x speedup.
Q2: When would you use cosine similarity vs Euclidean distance? ▼
Answer: Use cosine similarity for text/NLP embeddings (magnitude-independent, what transformers optimize for). Use Euclidean distance for geometric data where magnitude matters (images, coordinates). Cosine treats [1,0,0] and [10,0,0] as identical (same direction). Euclidean treats them differently (different magnitude).
Q3: How does HNSW indexing work at a high level? ▼
Answer: HNSW creates a multi-layered graph inspired by small-world networks. Vectors are organized in layers (L=0 is dense, L=high is sparse). To search: start at top layer, greedily navigate toward the query, then move down layers. Upper layers skip long distances quickly; lower layers find precise neighbors. Key parameters: M (connections per node) and ef_construction (search width during build).
Q4: What are the trade-offs of vector quantization? ▼
Answer: Benefits: 10-100x smaller vectors (less memory/storage/bandwidth), faster distance computation, enables massive scale on limited hardware. Costs: Slight accuracy loss (typically 1-5% recall loss with good quantization). Product quantization achieves 50-100x compression with <1% recall loss. Binary quantization gets 100x compression but larger accuracy loss.
Q5: How would you design a vector database to scale to 100 billion vectors? ▼
Answer: 1. Sharding: Split data across 1000+ nodes (100M vectors per node) 2. Replication: 3 replicas per shard for fault tolerance 3. Caching: L1 (node-local) + L2 (distributed) cache for hot vectors 4. Quantization: Use PQ or binary quantization to fit more in memory 5. Batch queries: Process 100-1000 queries in parallel 6. Monitoring: Track latency, recall, throughput, cache hit rates 7. Search optimization: Set ef parameters for 95-99% recall target (not 100%)
Q6: Explain the difference between IVF and HNSW indexing. ▼
Answer: IVF (Inverted File): Partitions space into clusters, searches only nearby clusters. Faster building but requires careful tuning (n_probe). HNSW: Creates layered graph structure. Slower building but better query performance and easier parameter tuning. HNSW is typically faster for queries; IVF better for very large datasets or batch operations.
Q7: What factors impact query latency in a vector database? ▼
Answer: 1. Index type and parameters (HNSW ef, IVF n_probe) 2. Vector dimensionality (higher = slower comparisons) 3. Number of vectors (larger index = more I/O) 4. Cache hit rate (cached vectors are 100x faster) 5. Network latency (for remote databases) 6. Concurrent queries (affects resource contention) 7. Filtering complexity (complex filters add latency) Typical latency: 1-10ms (memory), 10-100ms (disk), 50-500ms (network)
Q8: How do you choose between Pinecone, Qdrant, Weaviate, and ChromaDB? ▼
Answer: ChromaDB: Learning, prototyping, small projects (<1M vectors) Pinecone: Production SaaS, complex filtering needs, don't want to manage infrastructure Qdrant: Complex metadata filtering, cost-conscious, prefer self-hosting Weaviate: Knowledge graphs, hybrid search, enterprise/multi-tenant needs Decision factors: scale, filtering complexity, budget, ops preference, learning curve tolerance

Frequently Asked Questions

What is a vector database and why do I need it? ▼
A vector database stores high-dimensional numbers (embeddings) and finds similar ones efficiently. You need it for semantic search, RAG systems, recommendations, and any application where 'similarity' is important rather than exact matching.
Can I just use a regular SQL database for vectors? ▼
SQL databases are terrible at vector search. Scanning all rows is O(n), taking seconds or minutes for millions of records. Vector databases use specialized indexes (HNSW, IVF) making it O(log n) โ€” milliseconds instead.
What's the difference between embeddings and vectors? ▼
They're the same thing. 'Embedding' usually refers to the representation, 'vector' refers to the numerical form. An embedding is a vector. A text embedding is a vector representation of text created by a language model.
How much memory do vectors need? ▼
Each vector = dimensionality ร— bytes per dimension. 384-dim float32 = 1536 bytes. 1 million vectors = 1.5 GB. 1 billion vectors = 1.5 TB. With HNSW indexing overhead (pointers, graph): ~2-3x more. Use quantization if memory-constrained.
Can vector databases handle real-time updates? ▼
Yes, but with caveats. HNSW-based systems (Pinecone, Qdrant) handle insertions/deletions well without full reindexing. IVF-based systems (Milvus, Faiss) require periodic reindexing. For true real-time: expect latency impact during updates.
How do I measure if my vector database is working correctly? ▼
Use Recall@K: of the top-K results, how many are actual nearest neighbors? Target 95-99% recall. Measure latency at P50, P95, P99. Monitor cache hit rates. Compare results against brute-force baseline for small test sets.
What's the typical latency I should expect? ▼
Depends on setup: In-memory (1-10ms), SSD (10-100ms), Network (50-500ms), Brute-force (1-100 seconds). For production, target P99 latency < 100ms including everything (network, filtering, etc).
Can I combine vector search with keyword search? ▼
Yes! Weaviate calls this 'hybrid search'. Qdrant and others let you combine vector similarity with metadata filtering. For pure keyword + vector: use BM25 keyword search, then rerank using vectors.
How do I prevent hallucination in RAG systems using vector databases? ▼
1. Increase retrieval accuracy (higher Recall@K). 2. Retrieve more documents (top-20 instead of top-3). 3. Use diverse, high-quality source documents. 4. Implement answer verification (check if answer is in retrieved documents). 5. Use smaller, more focused vector indices.
What's the cost of running a vector database at scale? ▼
Pinecone: $200-1000/month (depending on vector count). Self-hosted (Qdrant/Milvus): $100-500/month (infrastructure only). pgvector: $50-200/month (Postgres). ChromaDB: Free (local) to $100/month (self-hosted). Calculate based on your query volume and data size.

Resources & Further Learning

Official Documentation

Research Papers

Tools & Libraries

  • Faiss โ€” Facebook's open-source library for similarity search
  • Annoy โ€” Approximate Nearest Neighbors in high-dimensional spaces
  • UMAP โ€” Dimensionality reduction for visualization
  • Numpy/SciPy โ€” Distance computations and linear algebra
  • sentence-transformers โ€” Creating embeddings easily

Related Topics

  • Embeddings & Representation Learning โ€” How text becomes vectors
  • Retrieval-Augmented Generation (RAG) โ€” Using vector search with LLMs
  • Information Retrieval โ€” Classical IR techniques (BM25, TF-IDF)
  • Approximate Nearest Neighbor Search โ€” The theoretical foundation
  • Dimensionality Reduction โ€” PCA, t-SNE, UMAP for visualization

Real-World Projects to Build

  • PDF/document search engine using embeddings
  • Semantic recommendation system for products
  • Question-answering system using RAG
  • Duplicate detection system for text/images
  • Search engine for your personal notes/knowledge base
  • Anomaly detection system using vector distance