Replace ChatGPT with a local AI: migration guide (2026)
Yes, for most common uses: writing, summarization, translation, code, and questions about your documents. The migration comes down to four steps: install Ollama, add an interface such as Open WebUI, choose a model that fits in your memory, then import your history ChatGPT. What you will not get locally is the capability level of the largest frontier models; keep them as a fallback for tasks where they make a difference.
Leaving ChatGPT doesn’t mean abandoning your habits. A chat interface, conversation history, preset assistants, file analysis, web search, and voice are all available locally. This guide maps each feature, tells you which model to choose based on your memory, and details how to bring your conversations back.
#What you gain, what you give up
Locally, your conversations and files stay on your machine: nothing passes through a third-party service, and the assistant works offline once the model has been downloaded. There is no subscription or message cap: the cost is the hardware and electricity. You also choose the model, its settings, and its instructions. In return, you lose automatic access to the largest cutting-edge models, which do not fit on a typical personal computer, as well as some of the turnkey integration: mobile apps, fluid voice, and agents connected to your accounts. The numerical comparison of performance and costs is in our guide “Local AI vs ChatGPT”.
| ChatGPT function | Local equivalent | What you need to know |
|---|---|---|
| Chat with history | Open WebUI, LM Studio, Jan, Msty | Open WebUI keeps conversations in a Docker volume |
| Import your old conversations | Import Chats from Open WebUI | The ChatGPT format is detected automatically |
| Custom GPTs | Custom Open WebUI models | Instructions, tools, and knowledge in one package |
| Memory | Open WebUI memory | Facts retained from one conversation to the next |
| Attachments, document analysis | Knowledge base (RAG) | The model’s context remains limited: see below |
| Web search | Web search for Open WebUI, SearXNG | Cited sources; depends on the configured engine |
| Voice | Speech recognition and synthesis from Open WebUI | Higher latency than a cloud service |
| Images | ComfyUI or Stable Diffusion | Model and GPU separate from the text model |
| State-of-the-art model | No complete equivalent | Keep an API or subscription as a fallback |
#Choose the model based on your memory
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
The rule is to choose the largest model that fits comfortably in Q4 within your VRAM or unified memory, including the context. The QuelLLM catalog gives the Q4 footprint of the weights alone for each model; the figures below are based on those values. Add the KV cache and a 20% margin; the context-window guide explains the calculation.
| Available memory | Possible model | Weights in Q4 | What to use it for |
|---|---|---|---|
| 8 GB | Qwen 3.5 9B | about 6 GB | Writing, summarization, common questions |
| 12 GB | Gemma 4 12B | approximately 7 GB | General-purpose assistant with context headroom |
| 16 GB | Mistral Small 3.2 24B or gpt-oss 20B | 14 to 13 GB | Reasoning and longer context |
| 24 GB | Qwen 3.8 27B | approximately 16 GB | Demanding daily use, coding |
| 32 GB | Qwen 3.6 35B-A3B (MoE) | about 21 GB | High throughput for its size, long contexts |
| 64 GB and more | Llama 3.3 70B or gpt-oss 120B | 40 to 70 GB | Quality close to current cloud models |
#Find a chat interface
- Open WebUI
- The web interface closest to ChatGPT: history, system prompts, attached files, web search, memory, and custom models. It connects to Ollama or any OpenAI-compatible API.
- LM Studio
- Desktop application that downloads the models and provides a chat: the fastest way to get started without a terminal.
- Jan
- Simple, privacy-focused desktop application with chat and model management.
- Msty
- Polished interface for Mac, Windows, and Linux, with multiple models side by side.
For a first installation, LM Studio or Jan are enough. Choose Open WebUI if several people need to use it, or if you want memory, custom models, or history import.
#Import your ChatGPT history
Open WebUI can read the ChatGPT export without prior conversion: the documentation specifies that the ChatGPT format is detected and converted automatically. Your conversations are added to the existing ones, and each import receives a new identifier: reimporting the same file creates duplicates. So perform the import une only once, after checking the file.
- 01Request the export from ChatGPTIn ChatGPT, open Settings, Data controls, then Export data and Request export. An email sends you a download link; the exact label depends on the interface language.
- 02Extract the archiveExtract the received file and locate conversations.json, which contains the history.
- 03Save Open WebUIBefore importing, export your existing Open WebUI conversations from Settings, Data Controls, Export Data, so you can roll back.
- 04ImportIn Open WebUI: username in the bottom left, Settings, Data Controls, Import Chats, then choose conversations.json. The label varies by language.
- 05ControlOpen a few older conversations: the messages should appear in order. The model shown for an imported conversation does not necessarily reflect the original model.
#Install the stack: Ollama, Open WebUI, a model
The following procedure assumes Docker is installed for Open WebUI. Without Docker, LM Studio or Jan replaces the whole setup in a single application.
- 01Install OllamaFollow your system’s installation guide. Ollama listens on http://localhost:11434 by default.
- 02Download a modelChoose based on the table above, for example ollama pull followed by the model name. Check its size before downloading it.
- 03Run Open WebUIThe official documentation provides a Docker command that mounts a volume to preserve your data and allows it to reach Ollama on the host machine.
- 04Adjust the contextOllama uses approximately 4,000 tokens with 24 GiB of VRAM, which truncates long documents. Increase it according to your available memory.
The volume option preserves your conversations across updates, and the secret key only needs to be set once: without a stable key, every container recreation disconnects users, the documentation explains. Replace your-secret-key with a generated value, for example using openssl rand -hex 32. The interface then responds on port 3000 of your machine.
#Find your use cases one by one
- Chat, write, translate
- Available as soon as a model is loaded. Quality depends on model size: a 9B is enough for an email, while 27B or larger is better for long, nuanced text.
- Analyze your documents
- The knowledge bases of Open WebUI, AnythingLLM, and PrivateGPT retrieve relevant passages instead of sending everything to the model. This is the right approach as soon as your documents exceed the context window.
- Coder
- Extensions such as Cline or Continue connect to Ollama. Ollama recommends at least 64 000 tokens of context for coding tools.
- Find your GPTs
- Recreate every GPT as a custom model: system instructions, tools, and knowledge bundled together. Copy your instructions from ChatGPT.
- Search the web
- Enable Open WebUI web search or connect SearXNG so the model cites sources.
#When to keep ChatGPT or use a cloud API as a fallback
Three situations justify not cutting it off completely: very long reasoning on a difficult problem, a task requiring a context of several hundred thousand tokens, or mobile use with continuous voice. A hybrid approach works well: local for anything confidential or repetitive, cloud for exceptions, never putting sensitive data there. Open WebUI lets you connect an external provider alongside Ollama and compare two models side by side, which helps you judge whether local is sufficient for your use case.
#Migrate safely: try it alongside the current setup first
Don’t cancel anything on day one. The safest approach is to run both tools side by side for a week or two, asking both the same question whenever the result matters. Note the cases where the local assistant disappoints you: they indicate whether you need a larger model, a larger context window, a better prompt, or, honestly, a use case to keep in the cloud.
- 01Days 1 and 2: install and testInstall the stack, load the model suited to your memory, and replay five typical conversations from your ChatGPT history.
- 02Days 3 to 7: double ChatGPTUse local models for everything routine; turn to ChatGPT only when the local answer is not enough, and note why.
- 03Week 2: import and recreateImport your history once, recreate your most-used GPTs as custom models, and enable web search if you need it.
- 04End of the period: decideCount the cases where you had to go back to the cloud. If that is rare, you can cancel the subscription; otherwise, keep it or switch to a usage-based API.
#Common disappointments after migration
- The context truncates your documents
- Ollama starts at around 4,000 tokens on a card with less than 24 GiB: a long PDF pasted into the chat is truncated without an obvious warning. Increase the window within the limits of your memory.
- A model that is too small or too compressed
- A 3B or 7B model at 2-bit quantization will make local AI seem mediocre. Increase the model size before drawing conclusions, and avoid quantization below 4 bits.
- A slow reasoning model
- Some models think for a long time before answering: the response arrives after several dozen seconds. Choose a model without reasoning for everyday conversation.
- Prompts written for ChatGPT
- A local model follows vague instructions less reliably. Specify the role, format, and response language in the system prompt.
#What still leaves your machine
Local does not mean sealed off. Open WebUI's web search sends your queries to the configured search engine; a model whose Ollama label ends in cloud runs on the provider's servers, like kimi-k3:cloud in the Ollama library; and connecting an external provider alongside Ollama, which Open WebUI allows, sends your messages to that provider. For real privacy, disable these options or reserve them for a separate profile, and check the list of installed models before processing sensitive documents.
Can you really replace ChatGPT with local AI?+
How do I retrieve my ChatGPT conversations?+
Which interface looks most like ChatGPT?+
Which model should you use with 16 GB of memory?+
What about my custom GPTs?+
Is it really free?+
#Go further
- Install Ollama in 5 minutes
- Open WebUI with Ollama: complete guide
- Local AI vs. ChatGPT: comparison
- Understanding the context window
- Local RAG with Ollama without coding
- SearXNG: web search for a local model
- Source: Open WebUI, import et conversation export
- Source: Open WebUI, features
- Source: Open WebUI, quick start
- Source: Ollama, context length
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.