Advanced 11 minStack

Milvus: the vector database for large volumes

Direct response

Milvus is an open-source vector database written in Go and C++, licensed under Apache 2.0 (more than 46,000 GitHub stars at the end of September 2026, version 3.0.2), designed for large-scale use: billions of vectors, with a distributed architecture that separates compute and storage. A lightweight variant, Milvus Lite, installs with pip install pymilvus and operates in a simple local file—useful for getting started, but rarely necessary for a modest local corpus.

Milvus is an open-source vector database designed for large-scale use: billions of vectors, distributed deployment, and a very broad choice of indexes. The project, written in Go and C++, had more than 46,000 GitHub stars at the end of September 2026, at version 3.0.2. On a single machine, it is also available in a lightweight version, Milvus Lite, letting you start small without switching tools later. The question is whether your corpus justifies this power—for many local projects, the answer is no, and that is useful information.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What Milvus is targeting

Most vector databases target team projects. Milvus targets industrial scale: separate storage and compute, native horizontal scaling on Kubernetes, and an index catalog that lets you finely balance accuracy, memory, and speed. The project claims it can process tens of thousands of queries over billions of vectors while keeping data up to date through real-time streaming updates. This is an architectural choice, not merely a collection of features.

This ambition has a direct downside: across a corpus of fifty thousand passages, you won’t see the benefit of that power, but you’ll immediately see the complexity. So the question isn’t “is this the best vector database?” but “will my corpus reach the size where these choices matter?”.

The project presents itself as the trusted foundation for AI developers building text and image search applications, retrieval-augmented generation, and recommendation systems, and claims to power many companies for use cases considered critical. This positioning clarifies its target audience: teams building a product intended to scale, not a personal document-search script for a few hundred PDF files.

#A component-based architecture

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

In a full deployment, Milvus is not a process but a collection of components: query nodes, data nodes, a coordination service, object storage, and a message log. Each part is sized independently—the project highlights the ability to scale query nodes separately for heavy read workloads and data nodes for heavy write workloads—which is exactly what you want at large scale and exactly what you can do without on a workstation. Stateless microservices on Kubernetes also enable fast recovery after an incident, and replica support further improves fault tolerance under load by loading data segments onto multiple query nodes. This separation of compute and storage is precisely what distinguishes a database designed for industrial scale from one designed for a single service: it has a permanent operating cost, even when traffic remains low, which few comparisons mention before deployment.

i
A lightweight mode is available: Milvus Lite
pip install -U pymilvus installe le SDK Python officiel ; pip install "pymilvus[milvus-lite]" ajoute la variante allégée. Il suffit ensuite d'instancier client = MilvusClient("milvus_demo.db") pour obtenir une base vectorielle locale dans un fichier, sans rien déployer. Le code écrit contre cette version reste valable si vous passez un jour au déploiement complet en changeant l'URI de connexion : c'est le principal argument pour commencer par Milvus plutôt que d'y migrer plus tard.

#The minimum in Python, with Milvus Lite

  1. 01
    Sign in
    from pymilvus import MilvusClient then client = MilvusClient("milvus_demo.db") opens a local database in a file; replacing the argument with a server URI switches to a full deployment without changing the rest of the code.
  2. 02
    Create a collection
    client.create_collection(collection_name="demo_collection", dimension=1024) — the dimension must match your embedding model’s exactly.
  3. 03
    Insert the data
    res = client.insert(collection_name="demo_collection", data=data), where data is a list of dictionaries, each containing a vector and its metadata.
  4. 04
    Search
    res = client.search(collection_name="demo_collection", data=vecteurs_requete, limit=5, output_fields=["text"]) returns the five closest passages with the requested fields.

The project highlights its integration with common AI tools—LangChain, LlamaIndex, OpenAI, and Hugging Face—which, according to its authors, makes it a vector store suited to retrieval-augmented generation. Milvus works with both open-source embedding models and embedding services, handling text, images, and video, and provides a utility (pymilvus[model]) to turn unstructured data into vectors without writing the model-call code yourself or separately managing each provider’s client library.

#Indexes, and which one to choose

Useful index families in practice
IndexTrade-offQuand
Exact (FLAT)Perfect accuracy, complete walkthroughUp to a few tens of thousands of vectors
Graph (HNSW)Fast queries, high memoryThe reasonable default below one million
Inverted file partitions (IVF)Fast build, settings to adjustLarge volumes, constrained memory
SCANNHigh-performance compressed vector searchAn alternative memory/speed trade-off to IVF-PQ
On disk (DiskANN)Capacity at the cost of latencyCorpus far exceeding available RAM
GPU (CAGRA)Hardware-accelerated building and searchVery large volumes, dedicated GPU available for indexing

For a local document corpus, keeping just one figure is usually enough: below one million vectors, the difference among the six index families is measured in milliseconds, not minutes, and the energy spent choosing the right setting would almost always be better invested elsewhere in the project.

The advice that saves the most time remains the same everywhere: start with exact search. Approximate indexes solve a scaling problem, and adopting them before you have that problem buys you parameters to tune and a recall metric to measure, in exchange for milliseconds nobody noticed. Milvus explicitly documents these six families—HNSW, IVF, FLAT, SCANN, DiskANN, and quantized variants—as each optimized for a different scenario, with GPU hardware support through NVIDIA's CAGRA for indexing very large volumes. The project adds metadata filtering and range search to these indexes, optimized at the same level as the vector search itself, rather than implemented as a separate layer that would slow the final result.

#What sets Milvus apart at scale

Beyond indexes, several features explain why large organizations choose Milvus over a simpler alternative. Its multi-tenancy can be configured with four possible isolation levels—database, collection, partition, or partition key—allowing a single cluster to serve anywhere from a few dozen to millions of tenants without sacrificing search performance or access-control granularity. Hot and cold storage keeps frequently accessed data in memory or on SSDs and moves rarely accessed data to slower, less expensive storage, reducing the bill without sacrificing the performance of critical tasks.

On the search side, Milvus natively supports full-text search with BM25 as well as learned sparse embeddings such as SPLADE and BGE-M3, in addition to dense-vector semantic search. Sparse and dense vectors can coexist in the same collection, with reranking functions to merge results from multiple queries—a hybrid search conceptually comparable to what Qdrant offers, but built here around BM25 rather than generic score fusion. On security, the project highlights mandatory authentication, TLS encryption for communications, and role-based access control (RBAC), three features expected as soon as a database serves multiple applications or teams, and much less critical for an instance listening only on localhost.

#Local use: what it entails

Resources
A full deployment requires several containers and several gigabytes of RAM before you even add your language model: the coordination service, query nodes, data nodes, object storage, and message log each run separately.
Vector memory
About 4 KB per 1,024-dimensional vector at full precision, excluding the index. This arithmetic—not the features—determines the architecture, and it remains true whether you use Milvus, Qdrant, or pgvector.
Deployment
Backups, upgrades, monitoring: a real distributed database requires a real operator. The Milvus ecosystem includes Attu, a graphical administration interface, and Birdwatcher for system debugging—two tools that nevertheless require knowledge of the underlying architecture.
The GPU remains dedicated to the model
Document encoding and inference are already competing for the card; vector search is mainly a CPU and memory task, unless GPU index construction is explicitly enabled through CAGRA for very large volumes.

#When Milvus is the right choice

The right question is never “which database has the most features?” but “which of these features will my project actually use?” Four-level multi-tenancy, hot and cold storage, RBAC, and BM25 hybrid search exist for organizations serving thousands of users with compliance requirements; on a personal workstation or for a small team’s internal tooling, these mechanisms remain inactive while their configuration complexity remains fully present. Milvus is justified when the project’s trajectory—not its current state—points toward these needs within a reasonable timeframe.

Choose honestly
SituationWhat works
Millions of vectors, continuous growthMilvus
Need to fine-tune the trade-off between memory and accuracyMilvus
Team corpus, filters, a few hundred thousand passagesA simpler dedicated database, such as Qdrant
PostgreSQL already installedPostgreSQL’s vector extension (pgvector)
Single-process applicationAn embedded database or a library such as FAISS

#FAQ

Is Milvus free?+
Yes, the engine is open source under the Apache 2.0 license and can be self-hosted at no licensing cost; its GitHub repository had more than 46,000 stars at the end of September 2026. A paid cloud offering (Zilliz Cloud) is available from the vendor in parallel and is not required for local use.
Do you need a GPU?+
No, not for routine search, which is a processor- and memory-intensive workload. Some index-building methods, especially CAGRA from NVIDIA, can use a GPU to accelerate indexing at very large scale, but that is not the typical case for a local corpus.
Can it be used on a single workstation?+
Yes, through Milvus Lite, its lightweight variant installed with pip install "pymilvus[milvus-lite]", which works in a local file without a server to deploy or a container to run. The full deployment, with its separate nodes, coordination service, and object storage, is clearly disproportionate for a simple personal workstation.
Milvus or Qdrant?+
Qdrant is easier to operate, installs with a single Docker command, and is more than sufficient at team scale, with equally rich metadata filtering. Milvus makes sense when targeting several million vectors, distributed scaling on Kubernetes, or fine-tuning indexes across six different families, including GPU support.
How much memory do you need for one million vectors?+
About 4 GB at full precision for 1,024-dimensional vectors, excluding the index structure, and much less with product quantization or an on-disk index such as DiskANN. Allow additional space for the index and metadata, as well as object-storage memory in a fully distributed deployment rather than a lightweight single-file setup.
Which tools support Milvus in production?+
The official ecosystem includes Attu for graphical administration, Birdwatcher for debugging, Prometheus and Grafana for monitoring, Milvus CDC for data synchronization, and connectors to Spark, Kafka, and Airbyte for building broader ingestion pipelines beyond what strictly local use requires.
Does Milvus support hybrid search like Qdrant?+
Yes, but with a different approach: Milvus natively combines dense and sparse vectors (BM25, SPLADE, BGE-M3) in the same collection, with reranking functions to merge the results. It is hybrid search built around conventional full-text search rather than generic score fusion like the RRF used by other vector databases, a choice that favors precision for queries containing exact terms.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.