Jan: the 100% open-source alternative to ChatGPT locale
Jan (jan.ai) is a free desktop application under the AGPL license that resembles ChatGPT but runs entirely on your machine. No account, no cloud, no hidden telemetry — just a binary to install and open-weight models to load. This Jan AI local tutorial covers installation, the first chat, the OpenAI-compatible API server, and an honest comparison with LM Studio and Msty.
#Why Jan?
Three things set Jan apart in the crowded desktop LLM client landscape. First, it's genuinely open source: code on GitHub (menloresearch/jan), AGPL-3.0 license, reproducible build. LM Studio and Msty are free but closed source. If transparency into the binary running continuously on your machine matters to you, Jan is the only one of the three that checks this box.
Second, the offline-first philosophy is radical. Jan makes no network calls at startup. No automatic update checks, no telemetry, no disguised “phone home.” Models download on demand from Hugging Face, period. You can use it on an air-gapped machine without breaking anything.
Third, the architecture is clean: Jan is an Electron frontend on top of Cortex, a separate inference engine (formerly Nitro). Cortex uses llama.cpp as its backend and exposes a REST API. So you can use the Cortex server without the Jan frontend, or connect other clients to Cortex. It's more modular than LM Studio, which is monolithic.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- System
- Windows 10/11 (x64), macOS 12+ (Intel or Apple Silicon), Linux (deb, AppImage, or rpm).
- RAM
- 8 GB is the strict minimum (2B–3B models), 16 GB is comfortable (8B–9B Q4 models), and 32 GB+ targets 24B and larger.
- Disk space
- Allow 5 GB for Jan + Cortex, then 2 to 40 GB per model depending on size. A dedicated folder on an SSD is ideal.
- GPU (optional but recommended)
- NVIDIA with CUDA for Windows/Linux, automatic Metal on Apple Silicon. Without a GPU, Jan runs on the CPU—slow but functional for small models.
#1. Installation
Download the installer from the official website. Jan presents itself as a 1-click installer—no CLI manipulation and no Python dependencies to manage.
- 01Launch the installerDouble-click the .exe (Windows), .dmg (macOS), or .AppImage (Linux). On Linux, make the AppImage executable with chmod +x Jan-*.AppImage if necessary.
- 02First launchJan creates its data folder (~/jan on macOS/Linux, %APPDATA%\Jan on Windows). This is where the models, conversations, and config will live. Keep this in mind if you want to put it on a secondary SSD.
- 03Check CortexAt startup, Jan launches the Cortex service in the background. If a firewall window opens, allow the local connection (Cortex listens on 127.0.0.1, not on the network).
- 04OnboardingJan suggests a starter model (typically a small Qwen or Granite). You can accept it to test in 2 minutes, or skip it to choose one yourself in the Hub.
#2. First model
Jan includes a model Hub that points to Hugging Face with prepared and tested GGUF versions. Open the “Hub” tab in the left sidebar to view the catalog.
To get started, three solid choices in French depending on your VRAM:
- Small setup (8 GB RAM, no GPU)
- Granite 4.2 3B Q4_K_M (≈2.2 GB). Very lightweight, enough for chatting, summarizing a short text, and brainstorming. Decent tokens/sec on CPU. Even lighter: Qwen 3.5 2B (≈1.9 GB).
- Standard configuration (16 GB RAM, 6–8 GB GPU)
- Qwen 3.5 9B Q4_K_M (≈6.6 GB). The 2026 8 GB choice: 256k context, vision, and an excellent FR/EN compromise for most everyday uses.
- Comfortable setup (32 GB RAM, 12 GB+ GPU)
- Gemma 4 12B (multimodal, ≈7.6 GB) or Mistral Small 24B in Q4_K_M (≈14 GB, good in French). Reasoning is clearly better than 8–9B models.
Click "Download" on the model page. Jan displays the progress and stores the GGUF file in its data folder. Once downloaded, the button becomes "Use"—one click loads the model into memory.
#3. Getting Started
The Jan interface follows ChatGPT conventions: a conversation sidebar on the left, a chat area in the center, and a contextual settings panel on the right. Three things to know to be productive:
- Threads
- Each conversation is an independent “Thread” with its own assigned model. You can open multiple threads in parallel (useful for comparing the response of a 7B and a 14B to the same question).
- Assistants
- Assistants tab: create personas with a system prompt, temperature, and preassigned model. Handy for keeping an “Assistant code” (Qwen3-Coder, temperature 0.2) and an “Assistant writing” (Mistral Small, temperature 0.8) separate.
- Generation parameters
- Right panel—temperature, top-p, top-k, max tokens, penalty frequency. Adjustable per thread. The context window (n_ctx) is configured at the model level under Settings → My Models.
#4. Extensions and Cortex
Jan has an extension-based architecture. Its main components are themselves internal extensions (Inference Cortex Extension, Model Extension, etc.), which eventually makes it possible to swap them out. On the user side, the third-party extension ecosystem is still younger than VS Code's, but a few are worth checking out:
- Alternative inference engines
- Beyond the default Cortex/llama.cpp, you can enable extensions that point to a remote OpenAI-compatible endpoint (Ollama, vLLM, or even OpenAI cloud if you accept leaving the local setup).
- Cortex standalone
- Cortex ships with Jan but can run on its own. cortex run qwen3.5:2b in the CLI starts a server on the default port without the Jan frontend. Useful for scripting or deploying on a headless server.
- Configurable Model Hub
- You can add custom model sources (a private HF repo, an internal mirror) by editing the Hub settings. Useful in business environments for distributing in-house fine-tuned models.
#5. API server mode
Jan's server mode exposes an OpenAI-compatible API on the local network. This transforms Jan from a simple chat client into a real backend for your scripts, n8n automations, or a VS Code copilot through Continue.dev.
- 01Enable the serverSettings → Local API Server. Check "Start Local API Server." Jan starts Cortex in server mode on port 1337 by default (adjustable).
- 02Choose the served modelAll downloaded models are exposed. Each has an identifier you can use in the "model" field of requests (visible in the ID column of My Models).
- 03Quick curl testVerify that the server responds to a standard /v1/chat/completions call. The syntax is identical to OpenAI’s.
#Jan vs LM Studio vs Msty
All three apps are desktop clients for local LLMs, with a clean UI and a local API server. The devil is in the details.
- Jan
- Open source (AGPL), strict offline-first, modular architecture with Cortex as a separate engine. Smaller model catalog than LM Studio. Ideal for those who want freedom and auditability.
- LM Studio
- Closed source but free. The richest Hugging Face catalog of the three (live search in HF). MCP support, native MLX on Mac, multi-token prediction. More features, but opaque.
- Msty
- Closed source, freemium (paid Aurum version). Best for RAG: document workspaces, split chat (compare two models side by side), and Knowledge Stacks. Less solid for pure chat, more oriented toward a “personal agent.”
#Troubleshooting
- The model won't load / OOM
- Check the available VRAM/RAM. In Settings → My Models → [model] → Edit Parameters, lower n_gpu_layers to offload fewer layers to the GPU. Or switch to a smaller quantization (Q4_K_S instead of Q4_K_M, or IQ3_M).
- Very low tokens/sec
- Confirm that the GPU is actually being used: on NVIDIA, run nvidia-smi during generation; you should see Jan or cortex-server consuming VRAM. On Apple Silicon, Metal is enabled by default with no configuration required.
- Cortex does not start (Windows)
- Visual C++ Redistributable missing. Install vc_redist x64 from Microsoft. Also check that no other process is using port 1337 (change the port in settings if needed).
- Linux: AppImage won't launch
- Install libfuse2 (sudo apt install libfuse2 on Ubuntu 22.04+). For debugging, launch the AppImage from a terminal to see the errors.
- Model download stuck at 99%
- Stop and restart. Jan resumes the download. If the problem persists, download the GGUF manually from Hugging Face and import it via Settings → My Models → Import.
- Answers in English despite a French prompt
- Model not strong enough for consistent bilingual output. Add an explicit system prompt, « Always respond in French », to the assistant, or switch to a strong French model (Mistral Small 24B, Qwen 3.5, Gemma 4).
#Go further
Depending on where you want to go now:
- Compare Jan with its direct competitors
- “Ollama vs LM Studio vs Jan vs GPT4All” provides the detailed comparison for choosing the right tool for your profile.
- Understanding Q4/Q5/Q8 quantizations
- “Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoffs—useful for making an informed choice in Jan’s Hub.
- Run RAG on your documents
- Jan does not have a robust native RAG solution. “Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” presents no-code alternatives that also connect to Jan’s API server.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.