FAISS: the library behind search vectorielle
FAISS (Facebook AI Similarity Search) is an open-source library under the MIT license, developed by Meta’s AI research group, that searches for the vectors closest to a given vector. It isn’t a database: no server, metadata filtering, or built-in persistence. It had more than 41,000 GitHub stars at the end of September 2026 and serves as the internal engine for several vector databases.
FAISS is a similarity-search library for dense vectors, developed primarily by Meta’s fundamental AI research group and used inside countless tools that never mention it. It is not a database: no server, metadata filtering, access control, or persistence. For a local single-process document pipeline, it is the lightest option that works—and it stops being the right tool as soon as access permissions or multiple writers enter the picture.
#A library, not a database
FAISS is often compared with vector databases as though they were alternatives. They aren't the same category. A vector database is a service with an API, storage, filtering, and permissions. FAISS is a component: written in C++ with full wrappers for Python and NumPy, you give it vectors, it builds an in-memory index, and it answers the question “which ones are closest to this one.” Several vector databases use it, or something similar, internally. The project is released under the MIT license, had more than 41,000 GitHub stars at the end of September 2026, and its authors state that some of its methods scale to billions of vectors in main memory on a single server.
The practical consequence is that choosing FAISS means taking responsibility for everything else: saving the index to disk, reloading it, keeping it consistent with your documents, linking a vector's position to the text it came from, and deciding what happens when two processes try to write at the same time. None of this is provided by default; it is an intentional architectural choice, not a project omission.
#The indexes that matter
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
| Index | How it searches | When to use it |
|---|---|---|
| Exact (Flat) | Compare against all vectors | Up to a few tens of thousands: exact results, no tuning, genuinely fast enough |
| IVF (inverted partitions) | Divides the space into clusters and explores only a few partitions | Starting at hundreds of thousands; requires a training phase on representative data |
| HNSW (neighbor graph) | Browses a graph of links | Plenty of RAM available or a modest corpus: fast and accurate, but no vector deletion |
| PQ / OPQ (product quantization) | Stores compressed vectors as codes of M bytes | When the index no longer fits in memory: precision drops, memory usage rises much more |
| RaBitQ (maximum compression) | Compresses to 1 bit per dimension plus a small overhead | The final memory-saving measure, with a random-rotation step to maintain good accuracy |
The time-saving advice: start with exact search. Approximate indexes exist to solve a scaling problem; adopting them before you have that problem means buying parameters to tune and recall to measure in exchange for milliseconds nobody noticed. At the scale of a local corpus, retrieval is almost never the slow step—the language model's generation dominates the response time perceived by the end user.
#The minimum to get started in Python
- 01Install the librarypip install faiss-cpu installe le paquet officiel PyPI (version 1.15.1 fin septembre 2026). Pour la variante GPU, le projet documente une installation via conda : conda install -c pytorch -c nvidia -c conda-forge faiss-gpu=1.15.1.
- 02Build an exact indexindex = faiss.IndexFlatL2(dimension) creates an exact-search index for vectors with the dimensionality of your embedding model.
- 03Add the vectorsindex.add(vecteurs) where vecteurs is a NumPy array of shape (n, dimension) containing 32-bit floats.
- 04Querydistances, indices = index.search(query, k) returns the k nearest neighbors and their distances; you must map indices to the original text chunks.
- 05Save and reloadfaiss.write_index(index, chemin) followed by faiss.read_index(chemin) persists the index to disk between runs, since FAISS does not do this itself.
#Choose an index based on corpus size
The project’s official wiki publishes precise guidelines, expressed as strings to pass to its index factory (index_factory). Below one million vectors, IVF_K is sufficient, with K chosen between 4×√N and 16×√N depending on the number of vectors N, and a training set of 30×K to 256×K vectors. Between 1 and 10 million, the recommended combination is IVF65536_HNSW32, which uses HNSW to speed up cluster assignment. Between 10 and 100 million, use IVF262144_HNSW32; beyond that, up to one billion, use IVF1048576_HNSW32—at this point, training becomes significantly slower and is generally performed on a GPU while the rest runs on the CPU.
| Corpus size | Recommended configuration |
|---|---|
| Less than 1 million | IVF_K (K between 4×√N and 16×√N) |
| 1 to 10 million | IVF65536_HNSW32 |
| 10 to 100 million | IVF262144_HNSW32 |
| 100 million to 1 billion | IVF1048576_HNSW32 |
For a local document corpus—from a few thousand to a few hundred thousand text chunks—these guidelines mainly confirm that we are far from the threshold where an approximate index becomes necessary: an exact index or, at worst, a simple IVF_K covers nearly all real-world cases, while configurations with several hundred thousand training entries remain beyond ordinary document use.
#Memory, with the numbers in hand
A 1,024-dimensional vector in 32-bit floating point takes up about 4 KB. One million vectors therefore require roughly 4 GB, before the index structure itself. For an HNSW index specifically, the official wiki gives the formula (d×4 + M×2×4) bytes per vector, where d is the dimension and M is the number of links per vector (between 4 and 64: more links mean greater precision and more memory). This arithmetic determines most architectures: that’s why compression exists, and why a machine that also hosts a language model has less headroom than you might assume.
On compression, product quantization (PQ) encodes each vector in M bytes, typically up to 64—beyond that, scalar quantization (SQ) is generally just as accurate and faster. When compression quality really matters, the official guide recommends adding an OPQ transformation before quantization: it first reduces the vector's dimensionality through a linear transformation that makes it easier to compress, then applies product quantization to the result. This adds another computation step during indexing, but reduces precision loss compared with direct product quantization at the same code size. RaBitQ, the maximum-compression option, gets down to about (d/8 + 8) bytes per vector by retaining only one bit per dimension, at the cost of a random rotation step needed to maintain reasonable accuracy; variants with multiple bits per dimension are available to recover some accuracy with slightly more storage.
#The missing feature that decides everything
Real questions include conditions: only this client's documents, only after this date, only what this person is authorized to read. FAISS has no concept of metadata. The usual workaround—retrieve more results than necessary and then filter in Python—is wrong in a specific way: if the top fifty all belong to another department, filtering returns nothing, and your assistant says it found no information instead of saying it found none that you are allowed to see.
Access-control filtering, in particular, must not be implemented downstream of retrieval. This is the strongest practical argument for a system that filters during search: a vector database, or vectors in PostgreSQL where it is a WHERE clause.
- pgvector: filter during retrieval, in SQL
- Qdrant: the dedicated service
- Milvus: the vector database for large-scale workloads
- The complete RAG pipeline
- The QuelLLM local RAG kit
- Source: official FAISS repository on GitHub
- Source: official index-selection guide
- Source: FAISS quickstart in Python
#When FAISS is the right choice
- A single-process application
- Who loads an index at startup and queries it: a desktop tool, batch processing, or a notebook.
- A fixed corpus
- Rebuilt on a schedule rather than updated continuously.
- No per-user filtering
- Or filtering so coarse that an index by category remains reasonable.
- Critical latency
- When the cost of a network round trip to a database is precisely what you want to eliminate, for example in an embedded tool without a guaranteed connection.
Outside these cases, the service you avoid installing generally costs less over time than the persistence, filtering, and concurrency code you eventually end up rewriting yourself as the project grows.
One last useful reference before deciding: several vector databases you may encounter elsewhere do not magically replace FAISS; they wrap it or draw on the same index families (IVF, HNSW, product quantization) behind a network API, managed persistence system, and filtering engine. Understanding FAISS therefore means understanding a large part of how vector databases themselves work—a useful detour even if the final project uses Qdrant or Milvus instead of FAISS directly, because the same memory and accuracy tradeoffs appear under different parameter names.
#FAQ
Is FAISS a vector database?+
Is FAISS free and open source?+
How many vectors can it handle?+
Can you filter results by metadata?+
FAISS or a Vector Database?+
Which index should you start with?+
Can HNSW replace IVF in every case?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.