Milvus: the vector database for large volumes
Milvus is an open-source vector database written in Go and C++, licensed under Apache 2.0 (more than 46,000 GitHub stars at the end of September 2026, version 3.0.2), designed for large-scale use: billions of vectors, with a distributed architecture that separates compute and storage. A lightweight variant, Milvus Lite, installs with pip install pymilvus and operates in a simple local file—useful for getting started, but rarely necessary for a modest local corpus.
Milvus is an open-source vector database designed for large-scale use: billions of vectors, distributed deployment, and a very broad choice of indexes. The project, written in Go and C++, had more than 46,000 GitHub stars at the end of September 2026, at version 3.0.2. On a single machine, it is also available in a lightweight version, Milvus Lite, letting you start small without switching tools later. The question is whether your corpus justifies this power—for many local projects, the answer is no, and that is useful information.
#What Milvus is targeting
Most vector databases target team projects. Milvus targets industrial scale: separate storage and compute, native horizontal scaling on Kubernetes, and an index catalog that lets you finely balance accuracy, memory, and speed. The project claims it can process tens of thousands of queries over billions of vectors while keeping data up to date through real-time streaming updates. This is an architectural choice, not merely a collection of features.
This ambition has a direct downside: across a corpus of fifty thousand passages, you won’t see the benefit of that power, but you’ll immediately see the complexity. So the question isn’t “is this the best vector database?” but “will my corpus reach the size where these choices matter?”.
The project presents itself as the trusted foundation for AI developers building text and image search applications, retrieval-augmented generation, and recommendation systems, and claims to power many companies for use cases considered critical. This positioning clarifies its target audience: teams building a product intended to scale, not a personal document-search script for a few hundred PDF files.
#A component-based architecture
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
In a full deployment, Milvus is not a process but a collection of components: query nodes, data nodes, a coordination service, object storage, and a message log. Each part is sized independently—the project highlights the ability to scale query nodes separately for heavy read workloads and data nodes for heavy write workloads—which is exactly what you want at large scale and exactly what you can do without on a workstation. Stateless microservices on Kubernetes also enable fast recovery after an incident, and replica support further improves fault tolerance under load by loading data segments onto multiple query nodes. This separation of compute and storage is precisely what distinguishes a database designed for industrial scale from one designed for a single service: it has a permanent operating cost, even when traffic remains low, which few comparisons mention before deployment.
#The minimum in Python, with Milvus Lite
- 01Sign infrom pymilvus import MilvusClient then client = MilvusClient("milvus_demo.db") opens a local database in a file; replacing the argument with a server URI switches to a full deployment without changing the rest of the code.
- 02Create a collectionclient.create_collection(collection_name="demo_collection", dimension=1024) — the dimension must match your embedding model’s exactly.
- 03Insert the datares = client.insert(collection_name="demo_collection", data=data), where data is a list of dictionaries, each containing a vector and its metadata.
- 04Searchres = client.search(collection_name="demo_collection", data=vecteurs_requete, limit=5, output_fields=["text"]) returns the five closest passages with the requested fields.
The project highlights its integration with common AI tools—LangChain, LlamaIndex, OpenAI, and Hugging Face—which, according to its authors, makes it a vector store suited to retrieval-augmented generation. Milvus works with both open-source embedding models and embedding services, handling text, images, and video, and provides a utility (pymilvus[model]) to turn unstructured data into vectors without writing the model-call code yourself or separately managing each provider’s client library.
#Indexes, and which one to choose
| Index | Trade-off | Quand |
|---|---|---|
| Exact (FLAT) | Perfect accuracy, complete walkthrough | Up to a few tens of thousands of vectors |
| Graph (HNSW) | Fast queries, high memory | The reasonable default below one million |
| Inverted file partitions (IVF) | Fast build, settings to adjust | Large volumes, constrained memory |
| SCANN | High-performance compressed vector search | An alternative memory/speed trade-off to IVF-PQ |
| On disk (DiskANN) | Capacity at the cost of latency | Corpus far exceeding available RAM |
| GPU (CAGRA) | Hardware-accelerated building and search | Very large volumes, dedicated GPU available for indexing |
For a local document corpus, keeping just one figure is usually enough: below one million vectors, the difference among the six index families is measured in milliseconds, not minutes, and the energy spent choosing the right setting would almost always be better invested elsewhere in the project.
The advice that saves the most time remains the same everywhere: start with exact search. Approximate indexes solve a scaling problem, and adopting them before you have that problem buys you parameters to tune and a recall metric to measure, in exchange for milliseconds nobody noticed. Milvus explicitly documents these six families—HNSW, IVF, FLAT, SCANN, DiskANN, and quantized variants—as each optimized for a different scenario, with GPU hardware support through NVIDIA's CAGRA for indexing very large volumes. The project adds metadata filtering and range search to these indexes, optimized at the same level as the vector search itself, rather than implemented as a separate layer that would slow the final result.
#What sets Milvus apart at scale
Beyond indexes, several features explain why large organizations choose Milvus over a simpler alternative. Its multi-tenancy can be configured with four possible isolation levels—database, collection, partition, or partition key—allowing a single cluster to serve anywhere from a few dozen to millions of tenants without sacrificing search performance or access-control granularity. Hot and cold storage keeps frequently accessed data in memory or on SSDs and moves rarely accessed data to slower, less expensive storage, reducing the bill without sacrificing the performance of critical tasks.
On the search side, Milvus natively supports full-text search with BM25 as well as learned sparse embeddings such as SPLADE and BGE-M3, in addition to dense-vector semantic search. Sparse and dense vectors can coexist in the same collection, with reranking functions to merge results from multiple queries—a hybrid search conceptually comparable to what Qdrant offers, but built here around BM25 rather than generic score fusion. On security, the project highlights mandatory authentication, TLS encryption for communications, and role-based access control (RBAC), three features expected as soon as a database serves multiple applications or teams, and much less critical for an instance listening only on localhost.
#Local use: what it entails
- Resources
- A full deployment requires several containers and several gigabytes of RAM before you even add your language model: the coordination service, query nodes, data nodes, object storage, and message log each run separately.
- Vector memory
- About 4 KB per 1,024-dimensional vector at full precision, excluding the index. This arithmetic—not the features—determines the architecture, and it remains true whether you use Milvus, Qdrant, or pgvector.
- Deployment
- Backups, upgrades, monitoring: a real distributed database requires a real operator. The Milvus ecosystem includes Attu, a graphical administration interface, and Birdwatcher for system debugging—two tools that nevertheless require knowledge of the underlying architecture.
- The GPU remains dedicated to the model
- Document encoding and inference are already competing for the card; vector search is mainly a CPU and memory task, unless GPU index construction is explicitly enabled through CAGRA for very large volumes.
#When Milvus is the right choice
The right question is never “which database has the most features?” but “which of these features will my project actually use?” Four-level multi-tenancy, hot and cold storage, RBAC, and BM25 hybrid search exist for organizations serving thousands of users with compliance requirements; on a personal workstation or for a small team’s internal tooling, these mechanisms remain inactive while their configuration complexity remains fully present. Milvus is justified when the project’s trajectory—not its current state—points toward these needs within a reasonable timeframe.
| Situation | What works |
|---|---|
| Millions of vectors, continuous growth | Milvus |
| Need to fine-tune the trade-off between memory and accuracy | Milvus |
| Team corpus, filters, a few hundred thousand passages | A simpler dedicated database, such as Qdrant |
| PostgreSQL already installed | PostgreSQL’s vector extension (pgvector) |
| Single-process application | An embedded database or a library such as FAISS |
- Qdrant: the dedicated service that’s simpler to operate
- pgvector: stay in PostgreSQL
- FAISS: the library rather than the service
- The complete RAG pipeline in Python
- The QuelLLM local RAG kit
- Source: official Milvus repository on GitHub
- Source: official Milvus Lite documentation
- Source: overview of the Milvus distributed architecture
#FAQ
Is Milvus free?+
Do you need a GPU?+
Can it be used on a single workstation?+
Milvus or Qdrant?+
How much memory do you need for one million vectors?+
Which tools support Milvus in production?+
Does Milvus support hybrid search like Qdrant?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.