Intermediate 11 minStack

Weaviate: hybrid search and multi-location

Direct response

Weaviate is an open-source vector database that can be installed in a container. It stands out in three ways: it can calculate vectors itself through modules (including Ollama, and therefore locally), it offers hybrid keyword-plus-vector search with an adjustable weight, and it isolates data by tenant. For strictly local RAG, three settings matter: a local vectorizer, telemetry disabled (it is enabled by default), and anonymous access disabled.

Weaviate is chosen for its features more than its simplicity: it is a full service, with a container, storage, and memory to size. This guide shows what it offers compared with ChromaDB or pgvector, how to connect it to Ollama so everything stays on your machine, how to configure hybrid search, what multitenancy is for, and what its memory costs.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What sets Weaviate apart from other vector databases

Weaviate is an open-source vector database whose code is published on GitHub. Unlike libraries such as FAISS, it is a server: it stores objects (text and properties), vectors, and the search index, and responds to queries through an API. Three features distinguish it from a simple vector store. The first is built-in encoding: you insert objects, a configured module transforms the text into vectors, and queries are written in natural language rather than as arrays of floating-point numbers. This choice eliminates a class of bugs (incompatible dimensions, a question encoded with a different model from the corpus), since the same module processes both sides.

The second is hybrid search, which combines keywords and vectors in a single query. The third is multi-tenancy: a single deployment can host multiple isolated datasets, one per customer, service, or user. The price of these features is a component that has to run, with its memory, backups, and updates, whereas ChromaDB in file mode is just a Python library.

#Keep everything local: three settings to check

The Local RAG Kit

Your documents, your AI: a reliable local RAG over your PDFs, notes and mail — nothing leaves your machine.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

An encoding module is a dependency with a location. Choosing a vectorizer hosted by a third party means sending each of your documents and each of your questions elsewhere; choosing the Ollama module keeps the computation on your hardware. Two other settings are less visible and deserve the same attention.

The vectorizer
Use text2vec-ollama, which calls your local Ollama instance; the documentation says it does not require an API key in this case. Watch the address: if Weaviate runs in a container and Ollama runs on the host machine, the documentation recommends host.docker.internal so the container can reach the host.
Telemetry
Weaviate's documentation indicates that it collects telemetry data by default: server version, operating system, modules used, number of objects and collections, sent every 24 hours; it specifies that none of your data's content is collected. To disable it, set DISABLE_TELEMETRY to true. In an installation that must remain hermetic, set it.
Anonymous access
The quick-start docker run command sets AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED to true, and the documentation strongly discourages anonymous access outside development or evaluation. As soon as port 8080 is reachable from anything other than your workstation, enable API-key authentication.
!
Encoding model version
Vectors produced by two different models aren’t comparable, even at the same dimensionality. Choose the model before the first bulk import: changing models means reimporting the entire collection. Choose a model that handles your corpus’s language correctly (bge-m3 for French, for example).

#Install Weaviate with Docker and Ollama

  1. 01
    Have Ollama and an embedding model
    Install Ollama and download the model with ollama pull bge-m3.
  2. 02
    Write the docker-compose.yml file
    It declares the Weaviate container, its data volume, activation of the Ollama module, disabled telemetry, and host access.
  3. 03
    Start and verify
    Run docker compose up -d, then test the address http://localhost:8080/v1/meta: the response lists the active modules.
  4. 04
    Create a collection linked to the module
    The collection indicates which model Ollama vectorizes which properties.
docker-compose.yml
services:
  weaviate:
    image: cr.weaviate.io/semitechnologies/weaviate:1.39.7
    ports:
      - "8080:8080"
      - "50051:50051"
    volumes:
      - weaviate_data:/var/lib/weaviate
    restart: on-failure:0
    extra_hosts:
      - "host.docker.internal:host-gateway"   # utile sous Linux pour joindre Ollama sur l'hôte
    environment:
      QUERY_DEFAULTS_LIMIT: 25
      AUTHENTICATION_ANONYMOUS_ACCESS_ENABLED: 'true'   # à fermer si le port est exposé
      PERSISTENCE_DATA_PATH: '/var/lib/weaviate'
      ENABLE_MODULES: 'text2vec-ollama'
      DISABLE_TELEMETRY: 'true'
      CLUSTER_HOSTNAME: 'node1'
volumes:
  weaviate_data:

Version 1.39.7 is the version of Weaviate’s official documentation at the time of writing; use the current version shown on the installation page. Ports 8080 (HTTP) and 50051 (gRPC) are the ones used in the documentation’s quick start.

#The schema: collections, properties, vectorizer, tenants

Weaviate concepts and their usefulness
ConceptWhat it isWhy it matters
CollectionA set of objects of the same type, with its schemaDistinct corpora remain separate, preserving precise retrieval
PropertiesTyped fields on every objectVerified filters (date, author, service) rather than instructions in the prompt
VectorizerThe module that encodes text and questionsSame model for indexing and querying
TenantAn isolated partition of the collection, with its own fragmentEach group sees only its own data
Vector indexThe search graph (HNSW), resident in memoryIt determines the speed and required memory
Create a vectorized collection with Ollama (Python client v4)
import weaviate
from weaviate.classes.config import Configure, Property, DataType

client = weaviate.connect_to_local()

client.collections.create(
    "Document",
    vector_config=[
        Configure.Vectors.text2vec_ollama(
            name="contenu_vecteur",
            source_properties=["contenu"],
            api_endpoint="http://host.docker.internal:11434",
            model="bge-m3",
        )
    ],
    properties=[
        Property(name="contenu", data_type=DataType.TEXT),
        Property(name="source", data_type=DataType.TEXT),
    ],
)

coll = client.collections.use("Document")
coll.data.insert_many([
    {"contenu": "Le délai de préavis est de trois mois.", "source": "contrat.pdf"},
    {"contenu": "Le loyer est révisé chaque année au 1er janvier.", "source": "bail.pdf"},
])
client.close()

The structure follows the official documentation: vector_config, a named vectorizer, source properties, and an Ollama access point. Weaviate vectorizes text-type properties by default, sorting them alphabetically and then concatenating them; source_properties lets you restrict the calculation to the relevant property and keep the filename out of the vector.

#Hybrid search in Weaviate

Semantic search finds passages with similar meaning and fails on exact strings: an invoice number, a part number, an error code, or a proper name. Weaviate's hybrid search combines the results of a vector search and a BM25F keyword search by merging the two result sets, with configurable weights and a configurable fusion method. The alpha parameter controls the balance: according to the documentation, 1 corresponds to pure vector search and 0 to pure keyword search. Without alpha, the effective weighting depends on your client: always set it explicitly.

Hybrid query
coll = client.collections.use("Document")
res = coll.query.hybrid(query="préavis contrat CDI", alpha=0.5, limit=5)
for o in res.objects:
    print(o.properties["source"], o.properties["contenu"][:80])

Since version 1.24, the default merge method has been relative-score merging; the alternative is rank merging. The principle and the choice between the two are explained in the guide on hybrid search. On a technical corpus, this is often the difference between a system you trust and one you abandon: pure-vector failures occur precisely on searches that users consider trivial.

#Multi-tenancy: one tenant per user group

Multi-tenancy partitions a collection into fragments, one per tenant. The documentation describes it as follows: each tenant is stored in a separate fragment, and one tenant's data is not visible to another. It is disabled by default and enabled in the collection definition with multi_tenancy_config. If multiple groups query the same system (SMB customers, company departments, family members), it is a structural answer to the question “can this person retrieve this document?”, much safer than a filter applied afterward, and infinitely safer than an instruction in the prompt.

Multi-tenant collection
from weaviate.classes.config import Configure
from weaviate.classes.tenants import Tenant

client.collections.create(
    "DocumentClient",
    multi_tenancy_config=Configure.multi_tenancy(enabled=True),
)
coll = client.collections.use("DocumentClient")
coll.tenants.create([Tenant(name="client_a"), Tenant(name="client_b")])

# Toute requête passe par un tenant : les autres restent invisibles
res = coll.with_tenant("client_a").query.hybrid(query="préavis", limit=3)

Tenants are lightweight: the documentation says you can have 50,000 or more active fragments per node. They have a state (ACTIVE, INACTIVE, OFFLOADED): an inactive tenant is stored on disk and uses no memory, making it possible to host many small datasets while keeping active only those in use. A tenant name accepts only alphanumeric characters, underscores, and hyphens.

#What it costs to run Weaviate

Weaviate is a real service: a container, persistent storage, and memory proportional to your vectors. The sizing documentation is clear about the constraint: the HNSW index must be stored in memory, memory determines the maximum dataset size, and it does not directly affect query speed. The documentation’s rule of thumb is to allow for twice the memory footprint of all vectors.

Memory calculation (documentation rule)
empreinte d'un vecteur = dimensions × 4 octets (float32)
mémoire estimée   = 2 × nombre de vecteurs × empreinte d'un vecteur

Exemple, bge-m3 (1 024 dimensions) :
  1 024 × 4 = 4 096 octets par vecteur
  100 000 passages → 2 × 100 000 × 4 096 ≈ 0,8 Go
  1 000 000 passages → 2 × 1 000 000 × 4 096 ≈ 8,2 Go

The 1,024 dimensions of bge-m3 are an assumption here that you should verify against your model's specifications. For a personal or small-business corpus (a few tens of thousands of passages), index memory usage is low; it becomes an issue at several million passages. Weaviate offers vector compression: the documentation recommends rotational quantization (RQ) and also mentions product quantization (PQ), binary quantization (BQ), and scalar quantization (SQ), at the cost of a slight loss of information. Add the local encoding module: Ollama runs an embeddings model on the same machine as your language model. On a single machine, decide which of the two gets the GPU, or accept that indexing and inference will compete with each other.

#Provide your own vectors, or migrate from ChromaDB

The encoding module is not required. Weaviate's documentation describes the “bring your own vectors” approach: instead of letting the database calculate embeddings, you provide ones you already have, whether customized or pre-generated. In the Python client, you then declare a named vector with Configure.Vectors.self_provided. This is the most cost-effective migration path from ChromaDB: you reread the documents and vectors already calculated (Chroma can return them with the include option), then send them to Weaviate without calling the embedding model again. Two checks prevent surprises: vector dimensions must be consistent across the entire collection, and queries must be encoded with the original model, since Weaviate will not do this for you.

When should you prefer built-in encoding? If you want queries to be written directly as text, adding documents to require no computation code, and the model's consistency to be guaranteed by the configuration. When should you prefer your own vectors? If you already have an embeddings pipeline, need to use a model that no module offers, or want to be able to change vector databases without recalculating everything.

#Weaviate or another vector database

Choose based on the situation
SituationChoice
Hybrid, with rich filters and a medium-to-large technical corpusWeaviate or Qdrant
Multiple isolated user groups on the same deploymentWeaviate (multi-location native)
PostgreSQL already in place, modest scalepgvector
Prototype or personal corpus, with no server to maintainChromaDB in file mode
Single process, fixed corpus, no filteringA library such as FAISS

If you are unsure, start with the simplest solution: ChromaDB for a prototype, then migrate to a service when a specific need arises (isolation, native hybrid search, volume). The choice is reversible as long as you keep the source documents and indexing script.

#Frequently asked questions about Weaviate

FAQ
Is Weaviate Free?+
The database is open source and self-hosted with no license cost; Weaviate separately offers a paid managed cloud service. Container-based self-hosting is the suitable option for a local installation. The actual costs are memory and administration: backups, updates, and disk-space monitoring.
Does Weaviate calculate embeddings itself?+
Yes, through modules called embedders, or it accepts vectors that you calculate yourself. With the text2vec-ollama module, computation is handled by your local Ollama instance, with no API key. An embedder hosted by a third party would send your documents and questions off the machine: avoid it for sensitive data.
Does Weaviate send data outside your environment?+
By default, it sends telemetry every 24 hours: version, system, modules, number of objects and collections, without your data contents according to the documentation. To disable it, set DISABLE_TELEMETRY to true in the configuration. For an installation that must remain isolated, do this during installation.
How do you tune alpha in hybrid search?+
Alpha equals 1 for pure vector search and 0 for pure keyword search. Start at 0.5, then measure recall in the first five results across 30 to 50 real questions: increase alpha if the questions use natural language; decrease it if they contain references and identifiers. Always set the value explicitly.
How much memory do you need for a million passages?+
With 1,024-dimensional float32 vectors, expect about 4 KB per vector, or 4 GB for one million passages, and about 8 GB according to the documentation’s rule of thumb, which doubles the footprint to account for the index. Compression (rotational quantization is recommended) substantially reduces this figure, at the cost of some accuracy.
Is multi-tenancy essential for personal use?+
No. A single user doesn't need partitions: one collection is enough. Multi-tenancy is useful as soon as multiple groups share a deployment and must not see one another's data (customers, services, households). It's disabled by default and enabled when the collection is created.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.