Weaviate: hybrid search and multi-location
Weaviate is an open-source vector database that can be installed in a container. It stands out in three ways: it can calculate vectors itself through modules (including Ollama, and therefore locally), it offers hybrid keyword-plus-vector search with an adjustable weight, and it isolates data by tenant. For strictly local RAG, three settings matter: a local vectorizer, telemetry disabled (it is enabled by default), and anonymous access disabled.
Weaviate is chosen for its features more than its simplicity: it is a full service, with a container, storage, and memory to size. This guide shows what it offers compared with ChromaDB or pgvector, how to connect it to Ollama so everything stays on your machine, how to configure hybrid search, what multitenancy is for, and what its memory costs.
#What sets Weaviate apart from other vector databases
Weaviate is an open-source vector database whose code is published on GitHub. Unlike libraries such as FAISS, it is a server: it stores objects (text and properties), vectors, and the search index, and responds to queries through an API. Three features distinguish it from a simple vector store. The first is built-in encoding: you insert objects, a configured module transforms the text into vectors, and queries are written in natural language rather than as arrays of floating-point numbers. This choice eliminates a class of bugs (incompatible dimensions, a question encoded with a different model from the corpus), since the same module processes both sides.
The second is hybrid search, which combines keywords and vectors in a single query. The third is multi-tenancy: a single deployment can host multiple isolated datasets, one per customer, service, or user. The price of these features is a component that has to run, with its memory, backups, and updates, whereas ChromaDB in file mode is just a Python library.
#Keep everything local: three settings to check
Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.
- Lifetime online access
- PDF + files
- Lifetime updates
An encoding module is a dependency with a location. Choosing a vectorizer hosted by a third party means sending each of your documents and each of your questions elsewhere; choosing the Ollama module keeps the computation on your hardware. Two other settings are less visible and deserve the same attention.
- The vectorizer
- Use text2vec-ollama, which calls your local Ollama instance; the documentation says it does not require an API key in this case. Watch the address: if Weaviate runs in a container and Ollama runs on the host machine, the documentation recommends host.docker.internal so the container can reach the host.
- Telemetry
- Weaviate's documentation indicates that it collects telemetry data by default: server version, operating system, modules used, number of objects and collections, sent every 24 hours; it specifies that none of your data's content is collected. To disable it, set DISABLE_TELEMETRY to true. In an installation that must remain hermetic, set it.
- Anonymous access
- The quick-start docker run command sets AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED to true, and the documentation strongly discourages anonymous access outside development or evaluation. As soon as port 8080 is reachable from anything other than your workstation, enable API-key authentication.
#Install Weaviate with Docker and Ollama
- 01Have Ollama and an embedding modelInstall Ollama and download the model with ollama pull bge-m3.
- 02Write the docker-compose.yml fileIt declares the Weaviate container, its data volume, activation of the Ollama module, disabled telemetry, and host access.
- 03Start and verifyRun docker compose up -d, then test the address http://localhost:8080/v1/meta: the response lists the active modules.
- 04Create a collection linked to the moduleThe collection indicates which model Ollama vectorizes which properties.
Version 1.39.7 is the version of Weaviate’s official documentation at the time of writing; use the current version shown on the installation page. Ports 8080 (HTTP) and 50051 (gRPC) are the ones used in the documentation’s quick start.
#The schema: collections, properties, vectorizer, tenants
| Concept | What it is | Why it matters |
|---|---|---|
| Collection | A set of objects of the same type, with its schema | Distinct corpora remain separate, preserving precise retrieval |
| Properties | Typed fields on every object | Verified filters (date, author, service) rather than instructions in the prompt |
| Vectorizer | The module that encodes text and questions | Same model for indexing and querying |
| Tenant | An isolated partition of the collection, with its own fragment | Each group sees only its own data |
| Vector index | The search graph (HNSW), resident in memory | It determines the speed and required memory |
The structure follows the official documentation: vector_config, a named vectorizer, source properties, and an Ollama access point. Weaviate vectorizes text-type properties by default, sorting them alphabetically and then concatenating them; source_properties lets you restrict the calculation to the relevant property and keep the filename out of the vector.
#Hybrid search in Weaviate
Semantic search finds passages with similar meaning and fails on exact strings: an invoice number, a part number, an error code, or a proper name. Weaviate's hybrid search combines the results of a vector search and a BM25F keyword search by merging the two result sets, with configurable weights and a configurable fusion method. The alpha parameter controls the balance: according to the documentation, 1 corresponds to pure vector search and 0 to pure keyword search. Without alpha, the effective weighting depends on your client: always set it explicitly.
Since version 1.24, the default merge method has been relative-score merging; the alternative is rank merging. The principle and the choice between the two are explained in the guide on hybrid search. On a technical corpus, this is often the difference between a system you trust and one you abandon: pure-vector failures occur precisely on searches that users consider trivial.
#Multi-tenancy: one tenant per user group
Multi-tenancy partitions a collection into fragments, one per tenant. The documentation describes it as follows: each tenant is stored in a separate fragment, and one tenant's data is not visible to another. It is disabled by default and enabled in the collection definition with multi_tenancy_config. If multiple groups query the same system (SMB customers, company departments, family members), it is a structural answer to the question “can this person retrieve this document?”, much safer than a filter applied afterward, and infinitely safer than an instruction in the prompt.
Tenants are lightweight: the documentation says you can have 50,000 or more active fragments per node. They have a state (ACTIVE, INACTIVE, OFFLOADED): an inactive tenant is stored on disk and uses no memory, making it possible to host many small datasets while keeping active only those in use. A tenant name accepts only alphanumeric characters, underscores, and hyphens.
#What it costs to run Weaviate
Weaviate is a real service: a container, persistent storage, and memory proportional to your vectors. The sizing documentation is clear about the constraint: the HNSW index must be stored in memory, memory determines the maximum dataset size, and it does not directly affect query speed. The documentation’s rule of thumb is to allow for twice the memory footprint of all vectors.
The 1,024 dimensions of bge-m3 are an assumption here that you should verify against your model's specifications. For a personal or small-business corpus (a few tens of thousands of passages), index memory usage is low; it becomes an issue at several million passages. Weaviate offers vector compression: the documentation recommends rotational quantization (RQ) and also mentions product quantization (PQ), binary quantization (BQ), and scalar quantization (SQ), at the cost of a slight loss of information. Add the local encoding module: Ollama runs an embeddings model on the same machine as your language model. On a single machine, decide which of the two gets the GPU, or accept that indexing and inference will compete with each other.
#Provide your own vectors, or migrate from ChromaDB
The encoding module is not required. Weaviate's documentation describes the “bring your own vectors” approach: instead of letting the database calculate embeddings, you provide ones you already have, whether customized or pre-generated. In the Python client, you then declare a named vector with Configure.Vectors.self_provided. This is the most cost-effective migration path from ChromaDB: you reread the documents and vectors already calculated (Chroma can return them with the include option), then send them to Weaviate without calling the embedding model again. Two checks prevent surprises: vector dimensions must be consistent across the entire collection, and queries must be encoded with the original model, since Weaviate will not do this for you.
When should you prefer built-in encoding? If you want queries to be written directly as text, adding documents to require no computation code, and the model's consistency to be guaranteed by the configuration. When should you prefer your own vectors? If you already have an embeddings pipeline, need to use a model that no module offers, or want to be able to change vector databases without recalculating everything.
#Weaviate or another vector database
| Situation | Choice |
|---|---|
| Hybrid, with rich filters and a medium-to-large technical corpus | Weaviate or Qdrant |
| Multiple isolated user groups on the same deployment | Weaviate (multi-location native) |
| PostgreSQL already in place, modest scale | pgvector |
| Prototype or personal corpus, with no server to maintain | ChromaDB in file mode |
| Single process, fixed corpus, no filtering | A library such as FAISS |
If you are unsure, start with the simplest solution: ChromaDB for a prototype, then migrate to a service when a specific need arises (isolation, native hybrid search, volume). The choice is reversible as long as you keep the source documents and indexing script.
#Frequently asked questions about Weaviate
Is Weaviate Free?+
Does Weaviate calculate embeddings itself?+
Does Weaviate send data outside your environment?+
How do you tune alpha in hybrid search?+
How much memory do you need for a million passages?+
Is multi-tenancy essential for personal use?+
- Hybrid search: BM25 + vector search
- Qdrant: the vector database for local RAG
- pgvector: vector search in PostgreSQL
- FAISS: the library behind vector search
- Local RAG with ChromaDB and Ollama
- Source: Weaviate, hybrid search
- Source: Weaviate, Ollama module
- Source: Weaviate, telemetry
- Source: Weaviate, resources and memory
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.