FAISS: Vector Search as a Library, Not a Server
FAISS is a similarity-search library you import, not a service you run. For a local RAG that fits in memory it is the fastest, leanest option — and the moment you need filtering, persistence or concurrency, it stops being the right shape.
Key takeaways
- FAISS is a library for similarity search over dense vectors, developed primarily at Meta's Fundamental AI Research group, MIT-licensed, with more than 41,000 GitHub stars as of late September 2026.
- It is not a database: no server, no metadata filtering, no access control, no durability by default. You call it from your process and manage everything around it.
- The project's own guidelines give exact thresholds by corpus size: under 1M vectors, a plain IVF index; 1M–10M,
IVF65536_HNSW32; 10M–100M,IVF262144_HNSW32; 100M–1B,IVF1048576_HNSW32. - Memory is the real constraint, and it is calculable. A 1,024-dimension float32 vector costs about 4 KB; an HNSW index adds
(d*4 + M*2*4)bytes per vector on top, per FAISS's own formula. - For a local RAG of a few thousand to a few hundred thousand chunks in one process, it is the leanest thing that works — you outgrow it the day you need per-user filtering, live updates from several processes, or persistence you do not want to hand-roll.
Library, not database
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- 30-day refund
Most discussions compare FAISS to vector databases as if they were alternatives. They are not the same category. A vector database is a service with an API, storage, filtering and access control. FAISS is a component, written in C++ with complete Python and NumPy wrappers: you hand it vectors, it builds an index in memory, and it answers "which of these are nearest to that." Several vector databases use it, or something like it, internally. The project is MIT-licensed, had more than 41,000 GitHub stars at the end of September 2026, and its own README states that some of its methods "can scale to billions of vectors in main memory on a single server."
The practical consequence is that choosing FAISS means accepting responsibility for the rest: saving the index to disk, reloading it, keeping it in sync with your documents, mapping vector positions back to the text they came from, and deciding what happens when two processes want to write at the same time. None of that ships by default — it is a deliberate design choice, not an oversight.
Getting running in Python
- Install it —
pip install faiss-cpupulls the official PyPI package (version 1.15.1 as of late September 2026); the project's own install guide documents a conda path for the GPU build:conda install -c pytorch -c nvidia -c conda-forge faiss-gpu=1.15.1. - Build an exact index —
index = faiss.IndexFlatL2(d), per the project's own getting-started guide, wheredis your embedding dimension. - Add vectors —
index.add(xb), a NumPy array of shape(n, d)in float32. - Search —
D, I = index.search(xq, k)returns distances and indices for theknearest neighbors; mappingIback to your source text is on you. - Persist —
write_index(index, "large.index")to save,read_index("large.index")to reload, since FAISS does not persist anything on its own.
The index types that matter
| Index | How it searches | Use it when |
|---|---|---|
| Flat (exact — IndexFlatL2 / IndexFlatIP) | Compares against every vector | Up to tens of thousands of vectors — exact results, no tuning, genuinely fast enough |
| IVF | Partitions the space, searches a few partitions | Hundreds of thousands and up; needs a training step on representative data |
| HNSW | Navigates a graph of neighbors | Plenty of RAM or a modest corpus: fast and accurate, but no vector removal |
| PQ / OPQ (product quantization) | Stores compressed vectors as M-byte codes | The index no longer fits in RAM; accuracy drops, memory drops far more |
| RaBitQ (maximum compression) | Compresses to roughly 1 bit per dimension plus overhead | The last resort for memory, with a random-rotation preprocessing step for accuracy |
The advice that saves the most time: start with exact search. The only indexes that guarantee exact results are IndexFlatL2 and IndexFlatIP, per FAISS's own documentation; they need no tuning, no training step, and no recall to measure. Approximate indexes exist to solve a scale problem, and adopting them before you have that problem buys tuning parameters and a training step in exchange for milliseconds nobody noticed. At local-RAG scale, retrieval is almost never the slow part; generation is.
Choosing an index by dataset size
FAISS's own wiki publishes concrete thresholds, expressed as strings you pass to its index factory. Below one million vectors, IVF_K_ is enough, with K chosen between 4×√N and 16×√N for N vectors, and a training set of 30×K to 256×K vectors. From one to ten million, the recommended combination is IVF65536_HNSW32, which uses HNSW to speed up cluster assignment during search. From ten to a hundred million, IVF262144_HNSW32; beyond that, up to a billion, IVF1048576_HNSW32 — at that scale, training slows down markedly and typically runs on GPU while everything else stays on CPU.
| Corpus size | Recommended configuration |
|---|---|
| Under 1 million | IVF_K (K between 4×√N and 16×√N) |
| 1–10 million | IVF65536_HNSW32 |
| 10–100 million | IVF262144_HNSW32 |
| 100 million–1 billion | IVF1048576_HNSW32 |
For a local document corpus — a few thousand to a few hundred thousand text chunks — these thresholds mostly confirm how far you are from needing an approximate index at all: exact search, or at worst a plain IVF_K, covers nearly every real case, and the training sets required at the higher tiers are far beyond typical local document use.
Memory, honestly
A vector of 1,024 dimensions in 32-bit floats occupies about 4 KB. A million of them is roughly 4 GB, before the index structure itself. For an HNSW index specifically, FAISS's own wiki gives the formula (d*4 + M*2*4) bytes per vector, where d is the dimension and M the number of links per vector (4 to 64: more links means more accuracy and more memory). That arithmetic decides most architectures: it is why compression exists, and why a machine that also holds a language model in its GPU memory has less room than people assume.
On compression: product quantization (PQ) encodes each vector into M-byte codes, typically 64 bytes or fewer — beyond that, scalar quantization (SQ) is usually about as accurate and faster. When compression quality really matters, the official guide recommends an OPQ transform first — a linear transformation to a reduced dimension that makes vectors easier to compress — before applying PQ on the result; it costs an extra step at indexing time but reduces accuracy loss compared to PQ alone at the same code size. RaBitQ, the maximum-compression option, brings that down to roughly (d/8 + 8) bytes per vector by keeping just one bit per dimension, at the cost of a random-rotation preprocessing step needed for decent accuracy; multi-bit variants trade some of that storage saving back for more precision.
Index in RAM, model in VRAM. FAISS indexes live in system RAM by default. GPU execution exists for large-scale workloads — the project's own documentation notes it "can accept input from either CPU or GPU memory," and a GPU index is a drop-in replacement for its CPU equivalent (GpuIndexFlatL2 for IndexFlatL2, for instance) — but competing for the same card as your language model is rarely a good trade in a local setup. Also worth knowing: HNSW only supports sequential adds (no add_with_ids without wrapping it in IDMap) and does not support removing vectors, unlike IVF.
The missing feature that decides most projects
Real questions carry conditions: only this client's documents, only after that date, only what this user may read. FAISS has no concept of metadata. The usual workaround — retrieve more results than needed, then filter in Python — is wrong in a specific way: if the top 50 all belong to another department, filtering leaves you with nothing, and your assistant reports that it found no information rather than that it found none you can see.
Permission filtering in particular should not be implemented downstream of retrieval. That is the strongest practical argument for a system that filters during the search — a vector database, or vectors inside PostgreSQL, where it is a WHERE clause.
When FAISS is the right call
- A single-process application that loads an index at startup and queries it — a desktop tool, a batch job, a notebook.
- A fixed corpus that is rebuilt on a schedule rather than updated continuously.
- No per-user filtering, or filtering so coarse that one index per category is reasonable.
- Latency-critical embedded search where the cost of a network hop to a database is the thing you are trying to remove.
Outside those cases, the service you avoid installing is usually cheaper, over time, than the persistence, filtering and concurrency code you end up rewriting as the project grows. Worth remembering too: several vector databases you will run into elsewhere do not replace FAISS by magic — they wrap it, or lean on the same index families (IVF, HNSW, product quantization) behind a network API, managed persistence and a filtering layer. Understanding FAISS is, to a large extent, understanding how vector databases work internally, even when the final project reaches for pgvector instead of FAISS directly.
Verdict
FAISS is excellent at exactly what it claims: fast nearest-neighbor search over vectors already in memory. In a local RAG it is the right choice for single-process, fixed-corpus, no-filtering situations — and the wrong one the moment several users, live updates or permissions enter the picture. Start with an exact index, measure before reaching for an approximate one, and move to a database — pgvector for a relational stack — when you catch yourself writing one. Readers building the rest of the stack can start from our local RAG toolkit.
Sources
- FAISS repository on GitHub — license, stars, README
- Guidelines to choose an index — official thresholds and memory formulas
- Getting started — official Python quickstart code
Frequently asked questions
Is FAISS a vector database?
No. It is a similarity-search library written in C++ with Python and NumPy wrappers. It has no server, no metadata filtering, no access control and no built-in durability — those are exactly what a vector database like Qdrant or Milvus adds on top of a comparable engine.
Is FAISS free?
Yes. It is MIT-licensed, a permissive license that allows commercial use at no royalty, and it is developed primarily at Meta's Fundamental AI Research group. There is no license fee, only the engineering time to integrate, operate and keep the library current inside your own stack over time.
How many vectors can it handle?
Millions, and according to the project's own authors potentially billions, with the right index and enough RAM. The binding constraint is memory: roughly 4 KB per 1,024-dimension vector in full precision, before the index structure, or far less with product quantization or RaBitQ for very large volumes.
Can I filter FAISS results by metadata?
Not during the search. You filter afterwards in your own code, which can leave you with nothing when the nearest results all fail the filter. For permission-based filtering, use a system that filters during retrieval itself, such as pgvector or a dedicated vector database with native payload filters.
FAISS or a vector database?
FAISS for a single process, a fixed corpus and no filtering. A database once you need several clients, live updates, persistence, or per-user access control enforced during the search rather than bolted on afterward in application code that has to be kept correct by hand.
Which index should I start with?
Exact search (IndexFlatL2 or IndexFlatIP, depending on your distance metric). No tuning, no training step, no recall to measure, and it stays fast well past the size of most local corpora — the official wiki only recommends an approximate index above roughly one million vectors.
Does GPU execution actually help?
Yes at large scale, and FAISS's own docs confirm GPU indexes are drop-in replacements for CPU ones with automatic memory copies. For a local corpus of a few hundred thousand vectors the gain is rarely noticeable, and it competes with your language model for the same card.
A current option for local AI: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395). Match memory to your model and software. A mini PC is a complete PC alternative; Mac/MLX and CUDA instructions require compatible hardware.
Amazon Check GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) price →As an Amazon Associate, BestLLMfor earns from qualifying purchases, at no extra cost to you. It does not influence our independent rankings.
Found an error or have feedback? Let us know — it helps everyone who reads this guide.