VectorOpsReport
Pink isometric server stack and a spiked node cluster linked by glowing pink lines and hexagons on a dark blue dotted grid, evoking distributed vector database nodes.
Comparison

Qdrant vs Weaviate vs Milvus: Recall, RAM, and Scale

This comparison covers filtered recall, memory and quantization options, cluster design, multi-tenancy, and scaling across the three engines.

By VectorOpsReport Editorial · · 5 min read

The qdrant vs weaviate vs milvus choice surfaces when a tenant_id filter drops recall@10 and triples p99 by shifting RAG retrieval from graph search to brute force. All three open source HNSW engines can look identical on vendor QPS charts. They differ in filtering, RAM cost, coordination beyond one node, and tenant limits. Sources are vendor documentation; vendor benchmarks are labeled.

What each one is, mechanically

Qdrant is a single Apache-2.0 Rust binary, v1.19.0. Its only dense index, HNSW, defaults to m: 16, ef_construct: 100. The filterable HNSW index adds graph edges from payload indexes (keyword, integer, float, bool, geo, datetime, text, uuid) to filter during traversal. Since v1.16.0, ACORN-style search explores neighbours-of-neighbours when direct ones are filtered. Vectors are always stored on disk with a cached or cold memory tier per collection.

Weaviate is BSD-3-Clause Go, v1.39.2, with maintained 1.38 and 1.37 lines. The vector index reference lists HNSW (defaults efConstruction: 128, maxConnections: 32, dynamic ef between 100 and 500), flat, dynamic (flat to HNSW after 10,000 objects), and HFresh for memory-constrained nodes. A model provider integration enables server-side vectorization through OpenAI, Cohere, Hugging Face, Ollama or one of a dozen other providers on insert.

Milvus is Apache-2.0 Go and C++. v3.0.0 released on 29 July 2026; maintained 2.6 includes 2.6.22, shipped on 4 August, and guarantees 2.6-to-3.0 compatibility and rollback. Disaggregated components are stateless proxies, one coordinator, streaming, query and data nodes, etcd, MinIO or S3, and the Woodpecker write-ahead log. The widest index menu includes FLAT, IVF_FLAT, IVF_SQ8, IVF_PQ, HNSW, HNSW_SQ, HNSW_PQ, HNSW_PRQ, SCANN, DiskANN, and four GPU indexes, including GPU_CAGRA. The CAGRA paper reports 33 to 77x HNSW’s large-batch throughput at 90 to 95% recall. HNSW defaults are M: 30 and efConstruction: 360.

The metric that matters

Use filtered QPS at target recall with RAM cost, not unfiltered QPS.

Define it on a golden set of N production queries with exact ground truth from a FLAT scan:

recall@k = mean over queries of |approx_topk ∩ exact_topk| / k
QPS@r    = sustained throughput at the ef where recall@k >= r

Vendor charts show unfiltered QPS; filters cause divergence. Weaviate pre-filters through its inverted index into an allow-list, walks HNSW with ACORN (default filterStrategy since v1.34), and uses brute force below flatSearchCutoff (default 40,000). Milvus restricts the search scope to matching entities before search, with iterative mode for complex expressions. Qdrant traverses its filter-aware graph and scans fully below full_scan_threshold. The ACORN paper reports 2 to 1,000x throughput at fixed recall over prior filtered-search approaches. Measure selectivity: a 1% tenant filter differs from a 40% region filter. See HNSW vs IVF tradeoffs for index-family differences.

Where they actually differ

ConcernQdrantWeaviateMilvus
Compressionscalar 4x, binary up to 32x, product up to 64x, TurboQuant up to 32x (docs)PQ (trained), SQ 4x, BQ 32x, RQ 8/4/1-bit (docs)IVF_SQ8, IVF_PQ, HNSW_SQ/PQ/PRQ, DiskANN on NVMe
Beyond RAMcold memory tier for vectors, index and payloadHFresh; tenant offload to S3DiskANN; query nodes page segments from object storage
ClusterRaft for topology; shard_number defaults to node count; replication_factor default 1 (docs)Raft for metadata, leaderless data replication with ONE/QUORUM/ALL (docs)Coordinator plus stateless workers on Kubernetes; etcd, object store, Woodpecker
Multi-tenancyone collection, is_tenant: true payload index; Cloud caps 1,000 collections (docs)one shard per tenant; docs cite ~170k active tenants on 9 n1-standard-8 nodes, bounded by the open-file limit (docs)database (64 default), collection (65,536 default), partition (1,024 per collection) or partition key routing to 16 partitions (docs)
Hybridsparse vectors, prefetch with RRF (v1.10) and DBSF (v1.11)BM25 plus vector, alpha default 0.75, relativeScoreFusion default since v1.24built-in BM25 function to SPARSE_FLOAT_VECTOR; 3.0 adds Block-Max WAND

Self-hosted Qdrant ignores post-creation replication_factor changes; add shard replicas manually. Qdrant Cloud reconciles automatically. The Milvus deployment guide defines Lite up to a few million vectors, Standalone up to 100 million, and Distributed from 100 million to tens of billions. See vector database memory sizing for RAM.

Wiring it up

One golden-set harness covers all three. An exact scan supplies ground truth for filtered recall@10 and p99.

import time, numpy as np
from qdrant_client import QdrantClient, models as qm
import weaviate
from weaviate.classes.query import Filter
from pymilvus import MilvusClient

queries = np.load("golden_queries.npy")          # (N, dim)
exact = np.load("golden_exact_top10.npy")        # (N, 10) ids from a FLAT scan
tenants = np.load("golden_tenants.npy")          # (N,) tenant per query
K = 10

def qdrant_search(c, q, t):
    r = c.query_points("docs", query=q.tolist(), limit=K,
        query_filter=qm.Filter(must=[qm.FieldCondition(
            key="tenant", match=qm.MatchValue(value=str(t)))]),
        search_params=qm.SearchParams(hnsw_ef=128))
    return [p.id for p in r.points]

def weaviate_search(coll, q, t):
    r = coll.query.near_vector(near_vector=q.tolist(), limit=K,
        filters=Filter.by_property("tenant").equal(str(t)))
    return [o.properties["doc_id"] for o in r.objects]

def milvus_search(c, q, t):
    r = c.search("docs", data=[q.tolist()], limit=K,
        filter=f'tenant == "{t}"', search_params={"params": {"ef": 128}})
    return [hit["id"] for hit in r[0]]

def evaluate(name, fn):
    hits, lat = [], []
    for q, gt, t in zip(queries, exact, tenants):
        t0 = time.perf_counter()
        ids = fn(q, t)
        lat.append(time.perf_counter() - t0)
        hits.append(len(set(ids) & set(gt.tolist())) / K)
    print(f"{name}: recall@{K}={np.mean(hits):.4f} "
          f"p50={np.percentile(lat,50)*1e3:.1f}ms p99={np.percentile(lat,99)*1e3:.1f}ms")

qc = QdrantClient("http://qdrant:6333")
wc = weaviate.connect_to_local(host="weaviate")
mc = MilvusClient("http://milvus:19530")
evaluate("qdrant",   lambda q, t: qdrant_search(qc, q, t))
evaluate("weaviate", lambda q, t: weaviate_search(wc.collections.get("Docs"), q, t))
evaluate("milvus",   lambda q, t: milvus_search(mc, q, t))

Sweep ef (Qdrant hnsw_ef, Weaviate ef, Milvus ef) from 32 to 512. Load-test the lowest value meeting target, not the default. Weaviate’s docs note diminishing recall gains above 512.

What you’ll see

Healthy recall-vs-ef curves flatten above target at similar ef, p99 rises roughly linearly with ef, and filtered and unfiltered curves stay close at every served selectivity.

An order-of-magnitude filtered p99 jump at one selectivity marks a brute-force cliff: Weaviate’s flatSearchCutoff, Qdrant’s full_scan_threshold, or a Milvus expression too complex for the standard path. Recall holding unfiltered but falling under a tight filter means a disconnected tenant graph; on Qdrant, a payload index created after ingestion usually omitted filter-aware edges. Week-over-week recall decay without config changes indicates embedding-model or corpus drift, covered in sentryml’s metrics taxonomy, not an engine issue. See low vector search recall: causes and fixes for triage order.

Caveats

Every cited head-to-head is a vendor benchmark. Qdrant’s benchmark page tests Qdrant, Weaviate, Milvus, Elasticsearch and Redis on an 8 vCPU, 32 GiB Azure D8s v3, caps each at 25 GB, covers dbpedia-openai-1M (1536d), deep-image-96 (10M), gist-960 and glove-100, and was last updated in 2024. VectorDBBench, sponsored by Milvus developer Zilliz, covers more than 30 engines across Cohere, OpenAI 500K and 5M, and LAION 100M, reporting QPS, recall, p99 and filtered-search cases. The only independent suite containing all three, ANN-Benchmarks, represents Qdrant, Weaviate and Milvus via Knowhere among 38 entries but tests single-query algorithms, not clustered load.

Quantization numbers are ceilings. Qdrant’s 40x binary-quantization speedup and Weaviate’s 98 to 99% 8-bit RQ recall are vendor results for specific embedding models. Below 8-bit requires rescore with oversampling on Qdrant or rescoring on Weaviate, reading full, uncompressed vectors from storage. GPU indexes target throughput, not latency: Milvus’s documentation says a GPU index “may not necessarily reduce latency compared to using a CPU index,” paying off under high request pressure or large query batches. GPU_CAGRA uses roughly 1.8x the raw vector data’s GPU-memory footprint.

Benchmarks omit tenant cardinality. Qdrant and Milvus vendors warn that one collection per tenant is expensive. Weaviate’s shard-per-tenant model scales further but remains bounded by the process open-file limit; offloaded tenants require S3. Qdrant’s shard_number cannot change without collection recreation and should be a multiple of expected node count. Retrieved chunks remain untrusted injection input, as aisec’s piece on indirect injection in RAG pipelines lays out.

Which one

Choose Qdrant for filtered search on one node or a handful with minimal operations; Weaviate for server-side embedding, out-of-box BM25 hybrid search, or tens of thousands of isolated tenants; Milvus past 100 million vectors, with Kubernetes and object storage, or for DiskANN or GPU indexes. Otherwise, use the harness and how to choose a vector database. For Qdrant and Milvus set against managed Pinecone in a RAG-pipeline context, RAGStackGuide’s Qdrant vs Milvus vs Pinecone comparison is the companion read.

Sources

  1. Qdrant docs: indexing (filterable HNSW, payload indexes, ACORN)
  2. Qdrant docs: quantization (scalar, binary, product, TurboQuant)
  3. Qdrant docs: distributed deployment (Raft, shards, replication_factor)
  4. Qdrant docs: multitenancy with payload partitioning
  5. Qdrant vector search benchmarks (vendor)
  6. Weaviate docs: vector index reference (HNSW defaults, flatSearchCutoff, filterStrategy)
  7. Weaviate docs: vector quantization (PQ, BQ, SQ, RQ)
  8. Weaviate docs: data structure and multi-tenancy
  9. Weaviate docs: replication architecture
  10. Milvus docs: deployment options (Lite, Standalone, Distributed)
  11. Milvus docs: in-memory index types
  12. Milvus docs: GPU index types
  13. Milvus docs: multi-tenancy strategies
  14. Milvus release notes (3.0.0)
  15. VectorDBBench, benchmark sponsored by Zilliz (vendor)
  16. ANN-Benchmarks (Aumüller, Bernhardsson, Faithfull), independent
  17. Malkov and Yashunin, HNSW (arXiv 1603.09320)
  18. Patel et al., ACORN: predicate-agnostic search over vector embeddings and structured data (arXiv 2403.04871)
  19. Ootomo et al., CAGRA: parallel graph construction and ANN search for GPUs (arXiv 2308.15136)
#qdrant#weaviate#milvus #vector-database #hnsw #comparison

Related