Qdrant: the vector database for RAG local
Qdrant is an open-source vector database written in Rust and released under the Apache 2.0 license (more than 34,000 GitHub stars as of late September 2026, version 1.19.1), started with a single Docker command. In a local document retrieval pipeline, it's the component that stores your document vectors, filters by their metadata during the search itself rather than afterward, and retrieves the passages closest to a question asked by the user.
Qdrant is a vector database written in Rust and released under the Apache 2.0 license. It installs with one command and comfortably handles corpora that in-memory libraries can no longer support. The project had more than 34,000 GitHub stars at the end of September 2026, at version 1.19.1. In a local document-retrieval pipeline, it stores your document vectors and retrieves the closest passages for each question. Here’s what it does well, and when a simpler solution is enough.
#What a vector database is for
An embedding model converts text into a list of numbers—a vector—such that two texts with similar meanings produce two nearby vectors. Finding the passages relevant to a question then amounts to searching for the vectors closest to the question’s vector. For a thousand passages, a simple calculation across the whole set is enough. For a million, you need an index structure, and that’s what a vector database does: build and maintain this structure, respond in a few milliseconds instead of comparing every vector one by one, and remain accurate as the corpus continues receiving new documents.
This step determines everything that follows: an excellent model that receives the wrong passages will answer poorly, and no prompt instruction can make up for faulty retrieval. That is why choosing the vector database—long treated as an interchangeable infrastructure detail—deserves the same care as choosing the language model itself.
#What Qdrant provides
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
The project describes itself as a “high-performance, large-scale” search engine and vector database, designed for extensive filtering—which distinguishes it from libraries that only search for the nearest neighbor without conditions. It is written in Rust, which its authors present as the reason for its speed and reliability under heavy load. In terms of operational robustness, the project also documents write-ahead logging that guarantees data persistence with update confirmation, even in the event of a power outage, along with metrics, telemetry, and audit logs for monitoring and debugging a production deployment.
- A standalone server
- One container, an HTTP and gRPC API, and official clients in Python, Go, Rust, JavaScript/TypeScript, .NET, and Java. It runs independently of your application, which can restart without ever losing the index that was built.
- Payload filtering
- Each vector carries metadata—author, date, service, document type—that we filter on during retrieval using should, must, and must_not clauses, not afterward. That's the difference between an enterprise-ready search and a demo.
- Vector quantization
- Compress vectors to reduce the memory footprint by a significant factor (4× in scalar form, up to 32× in binary form), at the cost of controlled, recoverable precision loss.
- Sparse vectors and multi-vectors
- Beyond standard dense vectors, Qdrant supports sparse vectors for full-text search and multivector objects, which are useful for late-interaction models such as ColBERT.
- Hybrid search
- Combine multiple vectors in a single query to benefit from both semantic understanding and keyword precision, with results merged using configurable strategies such as reciprocal rank fusion (RRF) or distribution-based score fusion (DBSF).
- Snapshots and recovery
- Save a collection and restore it elsewhere—what matters on the day reindexing would cost hours of GPU time.
- Distributed deployment
- Distribute a collection across multiple nodes through sharding and replication, with resizing without service interruption—a relevant approach beyond strictly local use, but useful to know if the project grows without rebuilding everything from scratch when a single machine is no longer enough.
#Get started locally
The shortest path is the official container, with a volume so the data survives restarts. An integrated web interface, described by the project as “a visual way to interact with your data and monitor the health of your deployment,” then lets you explore collections, manage data, and query the REST API without writing a single line of code—it’s the best diagnostic tool when a response is wrong: look at what was actually retrieved instead of guessing by rereading the retrieval code.
The Python client can also run with no server at all: QdrantClient(":memory:") for a throwaway test, or QdrantClient(path="path/to/db") for persistent local storage. This is valuable for a prototype or automated tests, with the same code later connecting to the server by changing one connection line. Since 2026, the project has documented a second integration path, Qdrant Edge: a lightweight version designed for resource-constrained devices that runs directly in the application process rather than in a client-server architecture, with optional synchronization to a full Qdrant server.
#The minimum in Python
- 01Sign infrom qdrant_client import QdrantClient then client = QdrantClient(url="http://localhost:6333") points to the container launched above.
- 02Create a collectionclient.create_collection(collection_name="docs", vectors_config=VectorParams(size=1024, distance=Distance.COSINE)) — the size must exactly match the dimension of your embedding model.
- 03Insert bullet pointsclient.upsert(collection_name="docs", points=[PointStruct(id=1, vector=[...], payload={"service": "support"})]) associates each vector with an identifier and filterable metadata.
- 04Queryclient.query_points(collection_name="docs", query=vecteur_question, limit=5).points returns the five closest passages, along with their score and payload.
#Metadata filtering, the feature you regret neglecting
In real-world use, a question is almost never asked across the entire corpus. You search documents from a department, after a given date, of a specific type, or accessible to the user asking the question. Qdrant applies these conditions during vector search, using a rich filter vocabulary — keyword matching, full-text search, numeric ranges, geolocation — combined with should, must, and must_not logical clauses. This consistently returns the right number of relevant results, whereas filtering afterward may leave you with nothing.
Access control deserves a separate mention: if several people query the same index, the permissions filter is what prevents a model from citing a document to someone who is not allowed to read it. No prompt instruction can replace this filter, and enforcing it at the database level rather than in application code prevents a new access point to the same collection from forgetting to apply it.
#Keep it in memory: vector quantization
| Vector storage | Approximate footprint | Effect on quality |
|---|---|---|
| Raw 32-bit floats | ≈ 4 GB | Reference |
| 8-bit scalar quantization | ≈ 1 GB (÷4, documented by Qdrant) | Usually negligible loss |
| Binary quantization | ≈ 128 MB (÷32, documented by Qdrant) | Real loss, to be recovered through verification of the top candidates |
The documented practice is to search compressed vectors, then rerank (rescore) the best candidates with the original vectors — Qdrant exposes an oversampling parameter to adjust this trade-off: at 2.4 with a limit of 100 results, 240 candidates are preselected from the quantized index before final reranking. You retain most of the accuracy while dividing memory usage by four or more, which is no luxury on a machine that also hosts a model. The official documentation also claims up to 40 times the speedup with binary quantization compared with the original vectors, a figure to verify on your own dataset rather than take for granted. The project summarizes all these compression options, combined with on-disk storage, by announcing a memory-use reduction of up to 97% — an order of magnitude that explains why quantization is presented as a core feature rather than a marginal setting.
#Qdrant or another one
The question to ask is not “what is the best vector database?” but “what does my project require today?” A test script needs nothing more than an in-memory library. A business application that already queries PostgreSQL benefits from adding a vector extension rather than another service. Qdrant becomes the right choice precisely when several of these needs come together: a service shared by multiple applications, fine-grained metadata filtering, a corpus that continues to grow, and a desire not to implement persistence or snapshots yourself. Reconsidering this choice regularly, rather than locking it in with the first prototype, helps avoid both over-architecting a modest project and undersizing one that has grown.
| Situation | What works |
|---|---|
| Prototype, a few thousand runs, a single script | An in-memory library or a local file is enough |
| You already have PostgreSQL and few vectors | A vector extension in your existing database |
| Shared service, fine-grained filtering, growing corpus | Qdrant |
| Embedded application, with no server to administer | The Python client’s local mode, or Qdrant Edge |
- A complete local RAG pipeline, from ingestion to response
- Choosing an embedding model for French
- Add a reranker to improve relevance
- pgvector: when PostgreSQL is enough
- The QuelLLM local RAG kit
- Source: Qdrant’s official GitHub repository
- Source: official quantization documentation
- Source: Qdrant quick start in Python
#FAQ
Is Qdrant free?+
Do you need a GPU for Qdrant?+
Qdrant or Chroma?+
How much RAM do you need?+
Can it be used without a server?+
Does quantization really reduce quality?+
Is Qdrant suitable for multiple clients in a single instance?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.