Intermediate 11 minStack

Qdrant: the vector database for RAG local

Direct response

Qdrant is an open-source vector database written in Rust and released under the Apache 2.0 license (more than 34,000 GitHub stars as of late September 2026, version 1.19.1), started with a single Docker command. In a local document retrieval pipeline, it's the component that stores your document vectors, filters by their metadata during the search itself rather than afterward, and retrieves the passages closest to a question asked by the user.

Qdrant is a vector database written in Rust and released under the Apache 2.0 license. It installs with one command and comfortably handles corpora that in-memory libraries can no longer support. The project had more than 34,000 GitHub stars at the end of September 2026, at version 1.19.1. In a local document-retrieval pipeline, it stores your document vectors and retrieves the closest passages for each question. Here’s what it does well, and when a simpler solution is enough.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What a vector database is for

An embedding model converts text into a list of numbers—a vector—such that two texts with similar meanings produce two nearby vectors. Finding the passages relevant to a question then amounts to searching for the vectors closest to the question’s vector. For a thousand passages, a simple calculation across the whole set is enough. For a million, you need an index structure, and that’s what a vector database does: build and maintain this structure, respond in a few milliseconds instead of comparing every vector one by one, and remain accurate as the corpus continues receiving new documents.

This step determines everything that follows: an excellent model that receives the wrong passages will answer poorly, and no prompt instruction can make up for faulty retrieval. That is why choosing the vector database—long treated as an interchangeable infrastructure detail—deserves the same care as choosing the language model itself.

#What Qdrant provides

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The project describes itself as a “high-performance, large-scale” search engine and vector database, designed for extensive filtering—which distinguishes it from libraries that only search for the nearest neighbor without conditions. It is written in Rust, which its authors present as the reason for its speed and reliability under heavy load. In terms of operational robustness, the project also documents write-ahead logging that guarantees data persistence with update confirmation, even in the event of a power outage, along with metrics, telemetry, and audit logs for monitoring and debugging a production deployment.

A standalone server
One container, an HTTP and gRPC API, and official clients in Python, Go, Rust, JavaScript/TypeScript, .NET, and Java. It runs independently of your application, which can restart without ever losing the index that was built.
Payload filtering
Each vector carries metadata—author, date, service, document type—that we filter on during retrieval using should, must, and must_not clauses, not afterward. That's the difference between an enterprise-ready search and a demo.
Vector quantization
Compress vectors to reduce the memory footprint by a significant factor (4× in scalar form, up to 32× in binary form), at the cost of controlled, recoverable precision loss.
Sparse vectors and multi-vectors
Beyond standard dense vectors, Qdrant supports sparse vectors for full-text search and multivector objects, which are useful for late-interaction models such as ColBERT.
Hybrid search
Combine multiple vectors in a single query to benefit from both semantic understanding and keyword precision, with results merged using configurable strategies such as reciprocal rank fusion (RRF) or distribution-based score fusion (DBSF).
Snapshots and recovery
Save a collection and restore it elsewhere—what matters on the day reindexing would cost hours of GPU time.
Distributed deployment
Distribute a collection across multiple nodes through sharding and replication, with resizing without service interruption—a relevant approach beyond strictly local use, but useful to know if the project grows without rebuilding everything from scratch when a single machine is no longer enough.

#Get started locally

The shortest path is the official container, with a volume so the data survives restarts. An integrated web interface, described by the project as “a visual way to interact with your data and monitor the health of your deployment,” then lets you explore collections, manage data, and query the REST API without writing a single line of code—it’s the best diagnostic tool when a response is wrong: look at what was actually retrieved instead of guessing by rereading the retrieval code.

Run Qdrant with persistence
docker run -p 6333:6333 -p 6334:6334 \
  -v "$(pwd)/qdrant_storage:/qdrant/storage" \
  qdrant/qdrant

The Python client can also run with no server at all: QdrantClient(":memory:") for a throwaway test, or QdrantClient(path="path/to/db") for persistent local storage. This is valuable for a prototype or automated tests, with the same code later connecting to the server by changing one connection line. Since 2026, the project has documented a second integration path, Qdrant Edge: a lightweight version designed for resource-constrained devices that runs directly in the application process rather than in a client-server architecture, with optional synchronization to a full Qdrant server.

#The minimum in Python

  1. 01
    Sign in
    from qdrant_client import QdrantClient then client = QdrantClient(url="http://localhost:6333") points to the container launched above.
  2. 02
    Create a collection
    client.create_collection(collection_name="docs", vectors_config=VectorParams(size=1024, distance=Distance.COSINE)) — the size must exactly match the dimension of your embedding model.
  3. 03
    Insert bullet points
    client.upsert(collection_name="docs", points=[PointStruct(id=1, vector=[...], payload={"service": "support"})]) associates each vector with an identifier and filterable metadata.
  4. 04
    Query
    client.query_points(collection_name="docs", query=vecteur_question, limit=5).points returns the five closest passages, along with their score and payload.
i
Vector dimensionality is not a detail
A collection is created for a given vector size and distance metric, which depend on the selected embedding model. Changing the embedding model requires recreating the collection and reindexing everything: this choice is made at the beginning, not along the way, and documenting the exact version of the model used avoids an unpleasant surprise when someone else takes over the project.

#Metadata filtering, the feature you regret neglecting

In real-world use, a question is almost never asked across the entire corpus. You search documents from a department, after a given date, of a specific type, or accessible to the user asking the question. Qdrant applies these conditions during vector search, using a rich filter vocabulary — keyword matching, full-text search, numeric ranges, geolocation — combined with should, must, and must_not logical clauses. This consistently returns the right number of relevant results, whereas filtering afterward may leave you with nothing.

Access control deserves a separate mention: if several people query the same index, the permissions filter is what prevents a model from citing a document to someone who is not allowed to read it. No prompt instruction can replace this filter, and enforcing it at the database level rather than in application code prevents a new access point to the same collection from forgetting to apply it.

#Keep it in memory: vector quantization

Order of magnitude for 1 million passes, 1,024-dimensional vectors
Vector storageApproximate footprintEffect on quality
Raw 32-bit floats≈ 4 GBReference
8-bit scalar quantization≈ 1 GB (÷4, documented by Qdrant)Usually negligible loss
Binary quantization≈ 128 MB (÷32, documented by Qdrant)Real loss, to be recovered through verification of the top candidates

The documented practice is to search compressed vectors, then rerank (rescore) the best candidates with the original vectors — Qdrant exposes an oversampling parameter to adjust this trade-off: at 2.4 with a limit of 100 results, 240 candidates are preselected from the quantized index before final reranking. You retain most of the accuracy while dividing memory usage by four or more, which is no luxury on a machine that also hosts a model. The official documentation also claims up to 40 times the speedup with binary quantization compared with the original vectors, a figure to verify on your own dataset rather than take for granted. The project summarizes all these compression options, combined with on-disk storage, by announcing a memory-use reduction of up to 97% — an order of magnitude that explains why quantization is presented as a core feature rather than a marginal setting.

#Qdrant or another one

The question to ask is not “what is the best vector database?” but “what does my project require today?” A test script needs nothing more than an in-memory library. A business application that already queries PostgreSQL benefits from adding a vector extension rather than another service. Qdrant becomes the right choice precisely when several of these needs come together: a service shared by multiple applications, fine-grained metadata filtering, a corpus that continues to grow, and a desire not to implement persistence or snapshots yourself. Reconsidering this choice regularly, rather than locking it in with the first prototype, helps avoid both over-architecting a modest project and undersizing one that has grown.

Choose based on the situation, not the trend
SituationWhat works
Prototype, a few thousand runs, a single scriptAn in-memory library or a local file is enough
You already have PostgreSQL and few vectorsA vector extension in your existing database
Shared service, fine-grained filtering, growing corpusQdrant
Embedded application, with no server to administerThe Python client’s local mode, or Qdrant Edge

#FAQ

Is Qdrant free?+
Yes, the engine is open source under the Apache 2.0 license and can be self-hosted with no licensing cost, with a GitHub repository that had more than 34,000 stars at the end of September 2026. A paid cloud offering also exists from the publisher; it is not required for local use.
Do you need a GPU for Qdrant?+
No for normal use: vector search, written in Rust, is a CPU and memory workload. Qdrant does document optional GPU support (NVIDIA and AMD) to accelerate index building on very large volumes — an option, not a prerequisite for a local corpus.
Qdrant or Chroma?+
Chroma is faster to set up for a Python prototype. Qdrant handles load better, offers finer-grained metadata filtering with a rich condition vocabulary, and is administered like a real service with snapshots, recovery, and built-in observability. The tipping point usually comes with multiple users or a few hundred thousand passages.
How much RAM do you need?+
It depends on the number of vectors and their dimension. One million 1,024-dimensional vectors take up approximately 4 GB raw, based on standard 32-bit floating-point arithmetic, and about four times less with the scalar quantization documented by Qdrant, or up to 32 times less with binary quantization. Add space for the index and metadata, which remain modest by comparison.
Can it be used without a server?+
Yes, the Python client can work in memory (":memory:") or on a local directory, which is suitable for prototypes and tests. The project also documents Qdrant Edge, a version that runs within the application process for resource-constrained environments, with synchronization to a full server when needed.
Does quantization really reduce quality?+
A little, but recoverable: the documented practice is to search compressed vectors, then reorder the best candidates using the original vectors (rescoring), with an adjustable oversampling parameter. 8-bit scalar quantization is generally the least risky; binary quantization is more aggressive and requires checking quality on your own corpus.
Is Qdrant suitable for multiple clients in a single instance?+
Yes, this is one of the use cases the project explicitly documents under the name multi-tenant: scalable data partitioning by customer or organization within the same collection, combined with payload filtering to isolate what each user can see. This is more economical, both operationally and in memory, than opening a separate collection for each customer once their number exceeds a few dozen.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.