GPT4All: first pas
GPT4All est une application de bureau gratuite (licence MIT, éditée par Nomic) qui fait tourner un modèle de langage sur votre ordinateur, sans GPU obligatoire et sans compte. Vous l'installez, téléchargez un modèle GGUF, et vous discutez. Sa dernière version, la 3.10.0, date de février 2025 : elle fonctionne, mais son catalogue est ancien et LocalDocs reste limité. Bon point d'entrée sur une machine modeste ; sinon, comparez avec Jan ou LM Studio.
GPT4All (often typed as “GPT for All”) is one of the oldest consumer-facing local chat applications. This guide explains what it does today, how to install it and load a first model, what LocalDocs and the local server are really like, and when it is better to choose something else.
#What GPT4All is and its status in 2026
GPT4All is developed by Nomic AI. The official repository describes it as a tool that runs language models privately on desktop and laptop computers, with no API call or mandatory GPU. The repository has about 77,000 stars, and the license is MIT. The application is built on llama.cpp and can load models in GGUF format.
The project’s status is the first thing to know. The latest published version is 3.10.0, dated February 25, 2025—more than a year and a half before this article. The historic gpt4all.io site now redirects to a Nomic page focused on its enterprise platform. The application remains usable, and nothing indicates that it will stop working, but you should not expect new features or an up-to-date inference engine.
#1. Installation
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
Installers are available on the official nomic-ai/gpt4all repository page (Windows, Windows ARM, macOS, Ubuntu). Download them only from this source: the page that opens at the historic gpt4all.io address is no longer the software's page. The prerequisites listed by the repository are: for Windows and Linux, a second-generation Intel Core i3 processor, AMD Bulldozer, or better; the Linux build is x86-64 only; for macOS, Monterey 12.6 or later, with better results on Apple Silicon chips in the M series.
| System | Installer | Prerequisites specified by the repository |
|---|---|---|
| Windows | Windows installer (and ARM version) | 2nd-generation Intel Core i3, AMD Bulldozer or better; ARM: Snapdragon, SQ1, SQ2 |
| macOS | macOS installer | Monterey 12.6 or later; Apple Silicon recommended |
| Linux | Ubuntu installer | x86-64 processor only, no ARM |
#2. First model
On first launch, the app suggests a default model. The documentation recommends starting with Llama 3, which you download using the model-add button. It notes that GPT4All connects to Hugging Face models through a llama.cpp engine, and that most use the .gguf extension.
- 01Open ModelsIn the left menu, click Models, then + Add Model to access the Explore Models page.
- 02Find a modelSearch among the models available online, or sort the results by popularity, downloads, or date via the gear icon.
- 03DownloadClick Download. The file is saved to your drive; no sign-up is required.
- 04Load and chatGo to Chats, click Load Default Model (Llama 3 or the model you just downloaded), and ask your first question.
The table below uses examples from the official documentation, along with the RAM indicated by the vendor. The “License” column matters: GPT4All lists models under very different terms, including one under a noncommercial license.
| Model | File | Required RAM | Size | License |
|---|---|---|---|---|
| Llama 3 Instruct | 4.66 GB | 8 GB | 8 billion, q4_0 | Meta Llama 3 License |
| Nous Hermes 2 Mistral DPO | 4.11 GB | 8 GB | 7 billion, q4_0 | Apache 2.0 |
| Phi-3 Mini Instruct | 2.18 GB | 4 GB | 4 billion, q4_0 | MIT |
| Mini Orca (Small) | 1.98 GB | 4 GB | 3 billion, q4_0 | CC-BY-NC-SA-4.0 (noncommercial use) |
Where are the models stored? You can set the download folder in the settings. By default, it is in AppData\Local\nomic.ai\GPT4All on Windows, Library/Application Support/nomic.ai/GPT4All on macOS, and .local/share/nomic.ai/GPT4All on Linux. A 4 to 5 GB model takes up the same amount of disk space: if your system drive is small, redirect this folder to a secondary drive before downloading several models.
#GPT4All in French
The application offers a language and regional-settings option for the interface. For responses, it all depends on the model: Llama 3 and models derived from Mistral understand and write French, with varying quality. Don’t rely on an English-language ranking: ask three questions in your field, in French, and evaluate the answer. If French is your primary use case, the Mistral and Qwen models in the QuelLLM catalog are better candidates, provided they load in the fixed engine.
#Lots of RAM on a processor: what it changes
A machine with 256 GB of RAM and no graphics card can load very large models, but generation speed depends on memory bandwidth, not capacity. The theoretical rule of thumb is: tokens per second are capped by bandwidth in GB/s divided by the model weight in GB. As an arithmetic illustration, 80 GB/s divided by 40 GB of weights gives at most 2 tokens per second. This is not a measurement, only a ceiling.
The only setting that can help is the processor thread count. In the advanced settings, GPT4All defaults it to 4; the documentation says that more threads can speed up responses. Increase it gradually, keeping some cores free for the system. If you have a graphics card, the device can be set to Auto, Metal, CPU, or GPU.
#3. The chat and settings that matter
The interface resembles a messaging app: conversation in the center, message box at the bottom, and model settings in the settings panel. Three default values deserve attention. The context length is 2,048 tokens, very short for a long document. The temperature is 0.7. And the number of layers placed in video memory is set to 32, which you should adjust according to your card.
#4. LocalDocs: built-in RAG and its limitations
LocalDocs lets you chat with your files without writing a single line of code. You create a collection by pointing it to a folder; the application splits your files into excerpts and indexes them with Nomic’s embedding models, running on your machine. In a chat, you enable the collection, and the model receives the excerpts semantically closest to your question in its prompt, with the sources displayed below the response.
- 01Create the collectionOpen LocalDocs, give the collection a name and choose the folder, then click Create Collection.
- 02Wait for indexingProgress is displayed; a green Ready indicator signals completion. You can already query the files that are ready.
- 03Enable it in a chatOpen LocalDocs using the button in the chat’s upper-right corner, then check the collection.
- 04Verify the sourcesClick Sources below the response to see which files were used.
#What LocalDocs does not do
Three documented limitations explain the disappointments. First, the file types indexed by default are .txt, .pdf, .md, and .rst: a Word file must be converted first. Next, each excerpt is 512 characters by default, and the application injects a maximum of three per query: the model therefore sees only about 1,500 characters from your documents for each question. Finally, this approach handles a targeted question well (“what does clause 4 say?”) and handles a summary of an entire case file poorly.
Advanced settings let you increase the size and number of excerpts, with this caveat from the documentation: they can improve response reliability but slow generation, and they consume context, which is already limited to 2,048 tokens by default. If you need to discuss a real document database, a tool such as AnythingLLM or Open WebUI is a better fit.
#5. The OpenAI-compatible local server
GPT4All includes an OpenAI-compatible API server, disabled by default. To enable it: Settings, Application, Advanced section, then check “Enable Local API Server.” It listens on port 4891, on localhost only (127.0.0.1), and accepts HTTP only, not HTTPS. The base address is http://localhost:4891/v1.
The available endpoints are /v1/models, /v1/models/nom-du-mod model, /v1/completions, and /v1/chat/completions. The model name in the request must match the one displayed in the application. One documented tip: LocalDocs can be enabled for the server chat from the interface (not through the API), and the references used are returned in the response, in choices[0].references.
#Privacy: three settings to know
By default, everything stays on your machine. Three options change that, and you need to know where they are. The Datalake option, which shares your interactions anonymously with the GPT4All community, is disabled by default. The “Use Nomic Embed API” option, which creates LocalDocs collections off-device using Nomic’s API, is disabled by default and requires a key. Finally, adding a “Model API” (a key from a remote provider) sends your prompts to that provider, as the documentation states unambiguously.
#Advantages, limitations, and alternatives
| Your situation | Is GPT4All a good fit? | An alternative worth considering |
|---|---|---|
| Modest machine, first try, no command line | Yes, that is its use case | Jan if you want a more active project |
| Chat with a few text files or PDFs | Yes, for targeted questions | AnythingLLM for a real document database |
| Very recent models (2025 and later) | Not guaranteed, fixed engine | LM Studio or Ollama, updated regularly |
| Shared server for a team | No: localhost only, HTTP | Ollama and Open WebUI |
- Strengths
- Short “download then chat” cycle, no account, LocalDocs without code, OpenAI-compatible server, works without a GPU.
- Weaknesses
- Engine frozen since February 2025, default context of 2,048 tokens, LocalDocs limited to three excerpts of 512 characters, outdated catalog.
- Ollama vs. LM Studio vs. Jan vs. GPT4All
- Jan: the open-source alternative to ChatGPT
- Get started with LM Studio
- AnythingLLM: production-ready RAG locally
- Open WebUI with Ollama: complete guide
- Source: GPT4All repository
- Source: GPT4All documentation, models
- Source: GPT4All documentation, settings
Is GPT4All free, and does it work offline?+
Is GPT4All still maintained in 2026?+
Which models should you use with GPT4All?+
Does GPT4All work in French?+
Does a machine with 256 GB of RAM make a difference?+
How can I query my documents with GPT4All?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.