Intermediate 11 minMigration

Replace ChatGPT with a local AI: migration guide (2026)

Direct response

Yes, for most common uses: writing, summarization, translation, code, and questions about your documents. The migration comes down to four steps: install Ollama, add an interface such as Open WebUI, choose a model that fits in your memory, then import your history ChatGPT. What you will not get locally is the capability level of the largest frontier models; keep them as a fallback for tasks where they make a difference.

Leaving ChatGPT doesn’t mean abandoning your habits. A chat interface, conversation history, preset assistants, file analysis, web search, and voice are all available locally. This guide maps each feature, tells you which model to choose based on your memory, and details how to bring your conversations back.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What you gain, what you give up

Locally, your conversations and files stay on your machine: nothing passes through a third-party service, and the assistant works offline once the model has been downloaded. There is no subscription or message cap: the cost is the hardware and electricity. You also choose the model, its settings, and its instructions. In return, you lose automatic access to the largest cutting-edge models, which do not fit on a typical personal computer, as well as some of the turnkey integration: mobile apps, fluid voice, and agents connected to your accounts. The numerical comparison of performance and costs is in our guide “Local AI vs ChatGPT”.

What you use in ChatGPT and its local equivalent
ChatGPT functionLocal equivalentWhat you need to know
Chat with historyOpen WebUI, LM Studio, Jan, MstyOpen WebUI keeps conversations in a Docker volume
Import your old conversationsImport Chats from Open WebUIThe ChatGPT format is detected automatically
Custom GPTsCustom Open WebUI modelsInstructions, tools, and knowledge in one package
MemoryOpen WebUI memoryFacts retained from one conversation to the next
Attachments, document analysisKnowledge base (RAG)The model’s context remains limited: see below
Web searchWeb search for Open WebUI, SearXNGCited sources; depends on the configured engine
VoiceSpeech recognition and synthesis from Open WebUIHigher latency than a cloud service
ImagesComfyUI or Stable DiffusionModel and GPU separate from the text model
State-of-the-art modelNo complete equivalentKeep an API or subscription as a fallback

#Choose the model based on your memory

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The rule is to choose the largest model that fits comfortably in Q4 within your VRAM or unified memory, including the context. The QuelLLM catalog gives the Q4 footprint of the weights alone for each model; the figures below are based on those values. Add the KV cache and a 20% margin; the context-window guide explains the calculation.

Recommended models by available memory (QuelLLM catalog, Q4, weights only)
Available memoryPossible modelWeights in Q4What to use it for
8 GBQwen 3.5 9Babout 6 GBWriting, summarization, common questions
12 GBGemma 4 12Bapproximately 7 GBGeneral-purpose assistant with context headroom
16 GBMistral Small 3.2 24B or gpt-oss 20B14 to 13 GBReasoning and longer context
24 GBQwen 3.8 27Bapproximately 16 GBDemanding daily use, coding
32 GBQwen 3.6 35B-A3B (MoE)about 21 GBHigh throughput for its size, long contexts
64 GB and moreLlama 3.3 70B or gpt-oss 120B40 to 70 GBQuality close to current cloud models
i
An MoE behaves differently
A mixture-of-experts model such as Qwen 3.6 35B-A3B requires the memory for its 35 billion parameters, but activates only 3 billion per token: it generates much faster than a dense model of the same size. See the guide to MoE models.

#Find a chat interface

Open WebUI
The web interface closest to ChatGPT: history, system prompts, attached files, web search, memory, and custom models. It connects to Ollama or any OpenAI-compatible API.
LM Studio
Desktop application that downloads the models and provides a chat: the fastest way to get started without a terminal.
Jan
Simple, privacy-focused desktop application with chat and model management.
Msty
Polished interface for Mac, Windows, and Linux, with multiple models side by side.

For a first installation, LM Studio or Jan are enough. Choose Open WebUI if several people need to use it, or if you want memory, custom models, or history import.

#Import your ChatGPT history

Open WebUI can read the ChatGPT export without prior conversion: the documentation specifies that the ChatGPT format is detected and converted automatically. Your conversations are added to the existing ones, and each import receives a new identifier: reimporting the same file creates duplicates. So perform the import une only once, after checking the file.

  1. 01
    Request the export from ChatGPT
    In ChatGPT, open Settings, Data controls, then Export data and Request export. An email sends you a download link; the exact label depends on the interface language.
  2. 02
    Extract the archive
    Extract the received file and locate conversations.json, which contains the history.
  3. 03
    Save Open WebUI
    Before importing, export your existing Open WebUI conversations from Settings, Data Controls, Export Data, so you can roll back.
  4. 04
    Import
    In Open WebUI: username in the bottom left, Settings, Data Controls, Import Chats, then choose conversations.json. The label varies by language.
  5. 05
    Control
    Open a few older conversations: the messages should appear in order. The model shown for an imported conversation does not necessarily reflect the original model.

#Install the stack: Ollama, Open WebUI, a model

The following procedure assumes Docker is installed for Open WebUI. Without Docker, LM Studio or Jan replaces the whole setup in a single application.

  1. 01
    Install Ollama
    Follow your system’s installation guide. Ollama listens on http://localhost:11434 by default.
  2. 02
    Download a model
    Choose based on the table above, for example ollama pull followed by the model name. Check its size before downloading it.
  3. 03
    Run Open WebUI
    The official documentation provides a Docker command that mounts a volume to preserve your data and allows it to reach Ollama on the host machine.
  4. 04
    Adjust the context
    Ollama uses approximately 4,000 tokens with 24 GiB of VRAM, which truncates long documents. Increase it according to your available memory.
Open WebUI with Docker (command from the official documentation)
docker run -d -p 3000:8080 --add-host=host.docker.internal:host-gateway -v open-webui:/app/backend/data -e WEBUI_SECRET_KEY=your-secret-key --name open-webui --restart always ghcr.io/open-webui/open-webui:main

The volume option preserves your conversations across updates, and the secret key only needs to be set once: without a stable key, every container recreation disconnects users, the documentation explains. Replace your-secret-key with a generated value, for example using openssl rand -hex 32. The interface then responds on port 3000 of your machine.

#Find your use cases one by one

Chat, write, translate
Available as soon as a model is loaded. Quality depends on model size: a 9B is enough for an email, while 27B or larger is better for long, nuanced text.
Analyze your documents
The knowledge bases of Open WebUI, AnythingLLM, and PrivateGPT retrieve relevant passages instead of sending everything to the model. This is the right approach as soon as your documents exceed the context window.
Coder
Extensions such as Cline or Continue connect to Ollama. Ollama recommends at least 64 000 tokens of context for coding tools.
Find your GPTs
Recreate every GPT as a custom model: system instructions, tools, and knowledge bundled together. Copy your instructions from ChatGPT.
Search the web
Enable Open WebUI web search or connect SearXNG so the model cites sources.

#When to keep ChatGPT or use a cloud API as a fallback

Three situations justify not cutting it off completely: very long reasoning on a difficult problem, a task requiring a context of several hundred thousand tokens, or mobile use with continuous voice. A hybrid approach works well: local for anything confidential or repetitive, cloud for exceptions, never putting sensitive data there. Open WebUI lets you connect an external provider alongside Ollama and compare two models side by side, which helps you judge whether local is sufficient for your use case.

#Migrate safely: try it alongside the current setup first

Don’t cancel anything on day one. The safest approach is to run both tools side by side for a week or two, asking both the same question whenever the result matters. Note the cases where the local assistant disappoints you: they indicate whether you need a larger model, a larger context window, a better prompt, or, honestly, a use case to keep in the cloud.

  1. 01
    Days 1 and 2: install and test
    Install the stack, load the model suited to your memory, and replay five typical conversations from your ChatGPT history.
  2. 02
    Days 3 to 7: double ChatGPT
    Use local models for everything routine; turn to ChatGPT only when the local answer is not enough, and note why.
  3. 03
    Week 2: import and recreate
    Import your history once, recreate your most-used GPTs as custom models, and enable web search if you need it.
  4. 04
    End of the period: decide
    Count the cases where you had to go back to the cloud. If that is rare, you can cancel the subscription; otherwise, keep it or switch to a usage-based API.

#Common disappointments after migration

The context truncates your documents
Ollama starts at around 4,000 tokens on a card with less than 24 GiB: a long PDF pasted into the chat is truncated without an obvious warning. Increase the window within the limits of your memory.
A model that is too small or too compressed
A 3B or 7B model at 2-bit quantization will make local AI seem mediocre. Increase the model size before drawing conclusions, and avoid quantization below 4 bits.
A slow reasoning model
Some models think for a long time before answering: the response arrives after several dozen seconds. Choose a model without reasoning for everyday conversation.
Prompts written for ChatGPT
A local model follows vague instructions less reliably. Specify the role, format, and response language in the system prompt.

#What still leaves your machine

Local does not mean sealed off. Open WebUI's web search sends your queries to the configured search engine; a model whose Ollama label ends in cloud runs on the provider's servers, like kimi-k3:cloud in the Ollama library; and connecting an external provider alongside Ollama, which Open WebUI allows, sends your messages to that provider. For real privacy, disable these options or reserve them for a separate profile, and check the list of installed models before processing sensitive documents.

FAQ
Can you really replace ChatGPT with local AI?+
For most common uses, yes: writing, summarization, translation, code, and questions about your documents work well with a model of 9 to 30 billion parameters. Very difficult reasoning tasks remain the domain of the largest cloud models. Test your own use cases before canceling a subscription.
How do I retrieve my ChatGPT conversations?+
In ChatGPT, request the export through Settings, Data Controls, Export Data. You will receive an archive by email containing conversations.json. Import this file into Open WebUI with Import chats: the format is detected automatically. Import it only once, or you will create duplicates.
Which interface looks most like ChatGPT?+
Open WebUI is the most complete: history, attachments, web search, memory, custom models, and direct import from the ChatGPT export. LM Studio and Jan are easier to install, with no Docker or terminal, but offer fewer team-work and customization features.
Which model should you use with 16 GB of memory?+
According to the QuelLLM catalog, Mistral Small 3.2 24B and gpt-oss 20B fit in Q4 with about 14 and 13 GB of weights. That leaves little room for context: limit the window, or choose a model with 9 to 12 billion parameters for long-context use.
What about my custom GPTs?+
They cannot be exported as-is. Copy their instructions and recreate each one as a custom model in Open WebUI, attaching your knowledge files to it. This takes a few minutes per assistant, and you can then choose the underlying model, which ChatGPT does not allow.
Is it really free?+
Most of the software mentioned is free and open source, and you don't need a subscription. You still pay for hardware, electricity, and setup time. If your current machine can only load a small model, the result will be more limited than ChatGPT.

#Go further

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.