Vector database selection becomes difficult when a RAG prototype meets production traffic. A database that looks fast on a small local dataset may deliver poor tail latency once you add metadata filters, concurrent users, frequent updates, multilingual embeddings, and limited cloud memory. Open source vector database benchmarks are useful only when they measure the workload your product will actually run.
This guide explains what to measure, how to compare Qdrant, Milvus, Weaviate, and FAISS, and how Indian AI teams can build a repeatable evaluation using realistic data and infrastructure.
What a useful benchmark should answer
A benchmark should help you answer four operational questions:
- Can the system return relevant results at your target recall?
- Does p95 or p99 latency remain acceptable during peak concurrency?
- How quickly can you ingest, update, delete, and rebuild data?
- What is the monthly cost at the required dataset size and availability level?
Avoid treating a single QPS number as a winner declaration. QPS depends on vector dimensions, index parameters, query distribution, hardware, filtering, replication, and whether results are served from memory or storage. Record the complete configuration with every result.
Teams building multilingual products should also test the actual embedding models and languages they support. A retrieval stack for Indic content may behave differently from one tested only on English benchmark datasets. Work on low-resource Indic natural language processing can also inform your decisions about corpus composition, tokenisation, and evaluation data.
Metrics to track
Recall at k
Recall@k compares approximate nearest-neighbour results with an exact-search ground truth. If Recall@10 is 0.95, the returned top ten contain 95% of the relevant neighbours measured by the chosen ground truth method. For RAG, test the k values your application uses—often 5, 10, or 20—and evaluate answer quality separately from retrieval recall.
Tail latency
Measure p50, p95, and p99 latency rather than average latency. P99 exposes queueing and contention that users experience during traffic spikes. Report latency for warm indexes and after restart, and separate database time from network, reranking, embedding, and application overhead.
Throughput under concurrency
Run queries at defined concurrency levels and identify the point where latency breaches your service-level objective. A system delivering high QPS at one client may perform poorly with dozens of simultaneous requests, writes, or filtered searches.
Ingestion and update performance
Measure initial bulk loading, index construction, incremental upserts, deletes, compaction, and recovery after failure. A knowledge base that changes daily needs different characteristics from a static archive. Include the time and resource cost of rebuilding indexes after changing HNSW or IVF parameters.
Memory, storage, and cost
Track resident memory, index size, raw vector size, payload size, replicas, CPU utilisation, and storage I/O. Convert these into cost per million vectors and cost per million queries using the same Mumbai or other India-region instance types. Include backups, observability, network transfer, and standby capacity.
How the main open-source options differ
Qdrant
Qdrant is a strong default for teams that want a focused vector database with a straightforward API, payload filtering, and efficient single-node or clustered deployments. Its Rust implementation and filtering capabilities make it attractive for production RAG where metadata constraints are common. Benchmark filtered and unfiltered searches separately; a result that is excellent on pure ANN search may not remain so when filters are selective.
Milvus
Milvus is designed for large collections and distributed operation. It is worth testing when you expect tens or hundreds of millions of vectors, multiple query nodes, high ingestion volume, or independent scaling of storage and compute. Distributed architecture introduces operational overhead, so compare not only peak throughput but also minimum viable cluster size, recovery time, and cost at your current scale.
Weaviate
Weaviate combines vector search with a developer-oriented data model and flexible metadata handling. It can be a good fit when teams want search, filtering, and application-level schema features in one system. Test hybrid search, filtered vector queries, updates, and module integrations rather than relying on raw ANN numbers.
FAISS
FAISS is an indexing and similarity-search library, not a complete distributed database. It is excellent for controlled experiments, local retrieval services, and custom systems where your team owns persistence, replication, APIs, deletes, and monitoring. It should not be compared directly with a full database unless those missing production components are included in the evaluation.
Managed services can provide a useful cost and operational baseline, but a self-hosted system should be compared at equivalent replication, availability, and support requirements—not just identical query latency.
Design an apples-to-apples test
Start with a representative corpus rather than a generic public dataset. Include the same vector dimensions, distance metric, payload fields, document length distribution, duplicate rate, and update pattern used in production. For Indian applications, include code-mixed queries, transliterated text, and regional-language content where relevant. Teams exploring open-source vision-language models for Indian languages should benchmark image or multimodal embeddings separately from text retrieval.
Use a fixed test matrix:
- Dataset sizes: 100,000; 1 million; 10 million; and your expected 12-month size.
- Recall targets: for example, 0.90, 0.95, and 0.99.
- Query types: pure vector, metadata-filtered vector, hybrid, and batch search.
- Concurrency: one client, normal peak, and overload conditions.
- Indexes: HNSW and IVF-based configurations where supported.
- Precision: FP32, reduced precision, or quantised vectors, measured with recall impact.
- Lifecycle events: restart, restore, compaction, reindexing, and rolling upgrade.
Use ANN-Benchmarks or an equivalent harness for algorithmic comparisons, then run an application-level test with your API, reranker, and LLM pipeline. Load generators such as Locust, k6, or JMeter can produce controlled concurrency. Pin software versions, container limits, CPU model, RAM, storage class, and network topology.
HNSW, IVF, and memory realities
HNSW commonly delivers strong recall and low latency but can consume substantial memory. Increase ef_search to improve recall and measure the corresponding latency and CPU cost. Index construction parameters such as M and ef_construction also affect build time and index size.
IVF methods divide the vector space into clusters and search a subset of them. They can reduce work and memory pressure, but recall depends on cluster count and probes. Do not compare HNSW and IVF at arbitrary settings: tune each to the same recall target, then compare latency, QPS, and cost.
Quantisation may reduce memory and improve cache efficiency, particularly at large scale. It can also reduce recall, especially for difficult or multilingual queries. Treat every compression setting as a separate operating point and validate downstream answer quality.
India-specific production checks
For an Indian startup, the cheapest benchmark winner is not necessarily the best production choice. Test in the region where users and application services run, and account for availability-zone traffic, managed disks, backups, and data-residency requirements. A lower-cost single node may be sensible for an early pilot, but production planning should include replica failure, restore time, and operational ownership.
Filter performance deserves special attention in fintech, healthcare, logistics, and commerce. Test combinations such as tenant ID, language, geography, document permissions, and freshness. Measure both selective and non-selective filters because execution plans can change significantly.
Keep retrieval evaluation tied to product outcomes. Build a labelled query set with relevant documents, then measure Recall@k, nDCG, citation accuracy, groundedness, and answer success. For broader guidance on production architecture, see building high-performance AI applications with open-source tools.
A practical decision rule
Choose the simplest system that meets your measured requirements:
- Choose Qdrant for a focused, efficient RAG service with strong filtering and manageable operations.
- Choose Milvus when distributed scale, ingestion volume, and horizontal query capacity justify a larger platform.
- Choose Weaviate when schema, hybrid retrieval, and developer ergonomics matter as much as raw ANN speed.
- Choose FAISS when you need a library for a custom retrieval service and are prepared to build the surrounding database capabilities.
Re-run benchmarks whenever you change the embedding model, corpus size, hardware, index parameters, or filtering logic. A benchmark is a decision tool, not a permanent ranking.
Benchmark checklist
Before selecting a database, confirm that you have:
- Matched vector dimensions, distance metric, and precision.
- Tested realistic metadata and tenant filters.
- Recorded p50, p95, p99, QPS, recall, memory, and storage.
- Included writes, deletes, restarts, compaction, and recovery.
- Compared equivalent replication and availability configurations.
- Calculated cost per useful query in your target India region.
- Validated retrieval quality on real multilingual and code-mixed queries.
Indian founders building retrieval infrastructure can also explore Indian open-source AI developer projects for implementation patterns and community context. If your team is developing an original AI infrastructure or application project, AI Grants India offers funding and ecosystem support through its grant programme.