All articles
Vector index tradeoffs explained from the documentation: HNSW against IVF, why recall drops, and how to size memory for an approximate search index.
-
What Is Reranking in Vector Search? A Practical Guide
Learn how reranking improves vector search relevance, measure nDCG and candidate recall, and instrument a PyTorch reranker without hiding latency costs.
-
Product Quantization Explained for Vector Search
Learn how product quantization compresses embeddings, how IVF-PQ changes recall, and how to evaluate a Faiss index with MLflow before deployment.
-
HNSW M Parameter Tuning: Recall, Memory, and Latency
The guide explains how M affects graph connectivity and memory, maps engine-specific settings, and shows a sweep for recall@10, p99, and bytes per vector.
-
How to Choose a Vector Database for Production
The guide compares pgvector, Qdrant, Milvus, Weaviate, Pinecone, and Elasticsearch by recall, filters, latency, memory, cost, and operational fit.
-
Cosine Similarity vs Dot Product: How Rankings Differ
Vector norms determine when cosine similarity and dot product rank results identically and when magnitude changes retrieval order.
-
How to Benchmark Recall at K for ANN Indexes
The guide explains exact ground truth, tie-safe recall@k, controlled efSearch sweeps, latency measurement, and how to interpret vendor benchmarks.
-
Hybrid Search: BM25, Vector Retrieval, and Score Fusion
This guide explains how BM25 and vector retrieval complement each other, why raw scores cannot be combined, and how RRF and alpha fusion work.
-
pgvector vs Pinecone: Cost, Recall, and Filtering
This comparison examines architecture, filtered recall, operational tradeoffs, and cost per query for self-hosted pgvector and managed Pinecone.
-
HNSW ef_search Parameter: Recall and Latency Tradeoffs
The HNSW ef_search parameter sets query beam width, balancing recall against latency across vector search engines and filtered queries.
-
Qdrant vs Weaviate vs Milvus: Recall, RAM, and Scale
This comparison covers filtered recall, memory and quantization options, cluster design, multi-tenancy, and scaling across the three engines.
-
HNSW vs IVF: Vector Index Tradeoffs Compared
HNSW vs IVF vector index tradeoffs: how graph and inverted-file designs compare on memory, build cost, recall tuning, updates, and filtering.
-
Low Vector Search Recall: Causes and Fixes
Why an approximate index returns the wrong neighbours: candidate lists too small, metric mismatch, filters, tombstones, and how to measure recall properly.
-
Vector Database Memory Sizing: RAM, Graph and Overhead
Size a vector index before you build it: bytes per embedding, HNSW graph overhead, what quantization actually saves, and what has to fit in RAM.
-
Vector Search Fundamentals: Embeddings, ANN and Recall
What an approximate nearest neighbor index does, how graph and cluster based indexes differ, and how quantization trades memory against recall.