Intermediate 11 minStack

FAISS: the library behind search vectorielle

Direct response

FAISS (Facebook AI Similarity Search) is an open-source library under the MIT license, developed by Meta’s AI research group, that searches for the vectors closest to a given vector. It isn’t a database: no server, metadata filtering, or built-in persistence. It had more than 41,000 GitHub stars at the end of September 2026 and serves as the internal engine for several vector databases.

FAISS is a similarity-search library for dense vectors, developed primarily by Meta’s fundamental AI research group and used inside countless tools that never mention it. It is not a database: no server, metadata filtering, access control, or persistence. For a local single-process document pipeline, it is the lightest option that works—and it stops being the right tool as soon as access permissions or multiple writers enter the picture.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#A library, not a database

FAISS is often compared with vector databases as though they were alternatives. They aren't the same category. A vector database is a service with an API, storage, filtering, and permissions. FAISS is a component: written in C++ with full wrappers for Python and NumPy, you give it vectors, it builds an in-memory index, and it answers the question “which ones are closest to this one.” Several vector databases use it, or something similar, internally. The project is released under the MIT license, had more than 41,000 GitHub stars at the end of September 2026, and its authors state that some of its methods scale to billions of vectors in main memory on a single server.

The practical consequence is that choosing FAISS means taking responsibility for everything else: saving the index to disk, reloading it, keeping it consistent with your documents, linking a vector's position to the text it came from, and deciding what happens when two processes try to write at the same time. None of this is provided by default; it is an intentional architectural choice, not a project omission.

#The indexes that matter

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Four families, three trade-offs
IndexHow it searchesWhen to use it
Exact (Flat)Compare against all vectorsUp to a few tens of thousands: exact results, no tuning, genuinely fast enough
IVF (inverted partitions)Divides the space into clusters and explores only a few partitionsStarting at hundreds of thousands; requires a training phase on representative data
HNSW (neighbor graph)Browses a graph of linksPlenty of RAM available or a modest corpus: fast and accurate, but no vector deletion
PQ / OPQ (product quantization)Stores compressed vectors as codes of M bytesWhen the index no longer fits in memory: precision drops, memory usage rises much more
RaBitQ (maximum compression)Compresses to 1 bit per dimension plus a small overheadThe final memory-saving measure, with a random-rotation step to maintain good accuracy

The time-saving advice: start with exact search. Approximate indexes exist to solve a scaling problem; adopting them before you have that problem means buying parameters to tune and recall to measure in exchange for milliseconds nobody noticed. At the scale of a local corpus, retrieval is almost never the slow step—the language model's generation dominates the response time perceived by the end user.

#The minimum to get started in Python

  1. 01
    Install the library
    pip install faiss-cpu installe le paquet officiel PyPI (version 1.15.1 fin septembre 2026). Pour la variante GPU, le projet documente une installation via conda : conda install -c pytorch -c nvidia -c conda-forge faiss-gpu=1.15.1.
  2. 02
    Build an exact index
    index = faiss.IndexFlatL2(dimension) creates an exact-search index for vectors with the dimensionality of your embedding model.
  3. 03
    Add the vectors
    index.add(vecteurs) where vecteurs is a NumPy array of shape (n, dimension) containing 32-bit floats.
  4. 04
    Query
    distances, indices = index.search(query, k) returns the k nearest neighbors and their distances; you must map indices to the original text chunks.
  5. 05
    Save and reload
    faiss.write_index(index, chemin) followed by faiss.read_index(chemin) persists the index to disk between runs, since FAISS does not do this itself.

#Choose an index based on corpus size

The project’s official wiki publishes precise guidelines, expressed as strings to pass to its index factory (index_factory). Below one million vectors, IVF_K is sufficient, with K chosen between 4×√N and 16×√N depending on the number of vectors N, and a training set of 30×K to 256×K vectors. Between 1 and 10 million, the recommended combination is IVF65536_HNSW32, which uses HNSW to speed up cluster assignment. Between 10 and 100 million, use IVF262144_HNSW32; beyond that, up to one billion, use IVF1048576_HNSW32—at this point, training becomes significantly slower and is generally performed on a GPU while the rest runs on the CPU.

Official benchmarks by corpus size
Corpus sizeRecommended configuration
Less than 1 millionIVF_K (K between 4×√N and 16×√N)
1 to 10 millionIVF65536_HNSW32
10 to 100 millionIVF262144_HNSW32
100 million to 1 billionIVF1048576_HNSW32

For a local document corpus—from a few thousand to a few hundred thousand text chunks—these guidelines mainly confirm that we are far from the threshold where an approximate index becomes necessary: an exact index or, at worst, a simple IVF_K covers nearly all real-world cases, while configurations with several hundred thousand training entries remain beyond ordinary document use.

#Memory, with the numbers in hand

A 1,024-dimensional vector in 32-bit floating point takes up about 4 KB. One million vectors therefore require roughly 4 GB, before the index structure itself. For an HNSW index specifically, the official wiki gives the formula (d×4 + M×2×4) bytes per vector, where d is the dimension and M is the number of links per vector (between 4 and 64: more links mean greater precision and more memory). This arithmetic determines most architectures: that’s why compression exists, and why a machine that also hosts a language model has less headroom than you might assume.

On compression, product quantization (PQ) encodes each vector in M bytes, typically up to 64—beyond that, scalar quantization (SQ) is generally just as accurate and faster. When compression quality really matters, the official guide recommends adding an OPQ transformation before quantization: it first reduces the vector's dimensionality through a linear transformation that makes it easier to compress, then applies product quantization to the result. This adds another computation step during indexing, but reduces precision loss compared with direct product quantization at the same code size. RaBitQ, the maximum-compression option, gets down to about (d/8 + 8) bytes per vector by retaining only one bit per dimension, at the cost of a random rotation step needed to maintain reasonable accuracy; variants with multiple bits per dimension are available to recover some accuracy with slightly more storage.

i
Index in RAM, model in VRAM
FAISS indexes live in system memory by default. GPU execution is available for very large volumes—the official documentation specifies that it accepts data from both CPU and GPU memory—but competing with your language model for the GPU is rarely a good trade-off in a local setup. Also note that an HNSW index only accepts sequential additions (no custom IDs without wrapping it in IDMap) and does not allow vectors to be deleted individually, unlike IVF.

#The missing feature that decides everything

Real questions include conditions: only this client's documents, only after this date, only what this person is authorized to read. FAISS has no concept of metadata. The usual workaround—retrieve more results than necessary and then filter in Python—is wrong in a specific way: if the top fifty all belong to another department, filtering returns nothing, and your assistant says it found no information instead of saying it found none that you are allowed to see.

Access-control filtering, in particular, must not be implemented downstream of retrieval. This is the strongest practical argument for a system that filters during search: a vector database, or vectors in PostgreSQL where it is a WHERE clause.

#When FAISS is the right choice

A single-process application
Who loads an index at startup and queries it: a desktop tool, batch processing, or a notebook.
A fixed corpus
Rebuilt on a schedule rather than updated continuously.
No per-user filtering
Or filtering so coarse that an index by category remains reasonable.
Critical latency
When the cost of a network round trip to a database is precisely what you want to eliminate, for example in an embedded tool without a guaranteed connection.

Outside these cases, the service you avoid installing generally costs less over time than the persistence, filtering, and concurrency code you eventually end up rewriting yourself as the project grows.

One last useful reference before deciding: several vector databases you may encounter elsewhere do not magically replace FAISS; they wrap it or draw on the same index families (IVF, HNSW, product quantization) behind a network API, managed persistence system, and filtering engine. Understanding FAISS therefore means understanding a large part of how vector databases themselves work—a useful detour even if the final project uses Qdrant or Milvus instead of FAISS directly, because the same memory and accuracy tradeoffs appear under different parameter names.

#FAQ

Is FAISS a vector database?+
No, it's a similarity-search library written in C++ with Python and NumPy bindings. No server, metadata filtering, access control, or built-in durability — that's precisely what a vector database such as Qdrant or Milvus adds on top of a comparable engine, often based on the same indexing principles.
Is FAISS free and open source?+
Yes. The project is released under the MIT license, a permissive license that allows commercial use without royalties, and is developed primarily by Meta's fundamental AI research group. No license fees to plan for—only the time required to integrate, operate, and keep the library up to date in your stack.
How many vectors can it handle?+
Millions, or even billions according to the project authors, with the right index and enough RAM. The constraint is RAM: about 4 KB per 1,024-dimensional vector at full precision, before the index structure, or significantly less with product quantization or RaBitQ for very large volumes.
Can you filter results by metadata?+
Not during retrieval. You filter afterward in your own code, which may leave you with nothing if all the best results fail the filter. For permission-based filtering, you need a system that filters during retrieval itself, such as pgvector or a dedicated vector database.
FAISS or a Vector Database?+
FAISS for a single process, a fixed corpus, and no filtering. Use a database as soon as you have multiple clients, continuous updates, persistence, or access rights that must be enforced during the search itself rather than afterward in your application code.
Which index should you start with?+
The exact index (IndexFlatL2 or IndexFlatIP depending on the distance). No tuning, no training phase, no recall to measure, and it is fast enough well beyond the size of most local corpora—the official wiki recommends an approximate index only beyond one million vectors.
Can HNSW replace IVF in every case?+
No. HNSW works well when RAM is plentiful or the corpus remains modest, but it only accepts sequential additions and does not support vector deletion. IVF is slower for pure query workloads but remains more flexible for a changing corpus and can be combined with HNSW beyond one million vectors. GPU execution directly replaces the equivalent CPU index (GpuIndexFlatL2 for IndexFlatL2, for example), but is rarely useful at the scale of a local corpus.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.