Beginner 10 minInterfaces

Jan: the 100% open-source alternative to ChatGPT locale

Jan (jan.ai) is a free desktop application under the AGPL license that resembles ChatGPT but runs entirely on your machine. No account, no cloud, no hidden telemetry — just a binary to install and open-weight models to load. This Jan AI local tutorial covers installation, the first chat, the OpenAI-compatible API server, and an honest comparison with LM Studio and Msty.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Jan?

Three things set Jan apart in the crowded desktop LLM client landscape. First, it's genuinely open source: code on GitHub (menloresearch/jan), AGPL-3.0 license, reproducible build. LM Studio and Msty are free but closed source. If transparency into the binary running continuously on your machine matters to you, Jan is the only one of the three that checks this box.

Second, the offline-first philosophy is radical. Jan makes no network calls at startup. No automatic update checks, no telemetry, no disguised “phone home.” Models download on demand from Hugging Face, period. You can use it on an air-gapped machine without breaking anything.

Third, the architecture is clean: Jan is an Electron frontend on top of Cortex, a separate inference engine (formerly Nitro). Cortex uses llama.cpp as its backend and exposes a REST API. So you can use the Cortex server without the Jan frontend, or connect other clients to Cortex. It's more modular than LM Studio, which is monolithic.

i
In one sentence
Jan = a polished Electron UI + Cortex (llama.cpp engine) + a Hugging Face model store. Think of a "ChatGPT desktop" that never leaves your Wi-Fi.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
System
Windows 10/11 (x64), macOS 12+ (Intel or Apple Silicon), Linux (deb, AppImage, or rpm).
RAM
8 GB is the strict minimum (2B–3B models), 16 GB is comfortable (8B–9B Q4 models), and 32 GB+ targets 24B and larger.
Disk space
Allow 5 GB for Jan + Cortex, then 2 to 40 GB per model depending on size. A dedicated folder on an SSD is ideal.
GPU (optional but recommended)
NVIDIA with CUDA for Windows/Linux, automatic Metal on Apple Silicon. Without a GPU, Jan runs on the CPU—slow but functional for small models.
→
The right VRAM benchmark
Q4_K_M is Jan’s default quantization. At this format: 3B≈2 GB VRAM · 7B≈5 GB · 14B≈9 GB · 32B≈19 GB · 70B≈40 GB. Above your VRAM capacity, Jan offloads to RAM (a sharp drop in tokens/sec).

#1. Installation

Download the installer from the official website. Jan presents itself as a 1-click installer—no CLI manipulation and no Python dependencies to manage.

Official website
https://jan.ai
!
Still from jan.ai or the official GitHub
The official repo is github.com/menloresearch/jan (formerly janhq/jan). Download Jan only from the official website or signed GitHub releases. Beware of precompiled forks distributed on third-party sites.
  1. 01
    Launch the installer
    Double-click the .exe (Windows), .dmg (macOS), or .AppImage (Linux). On Linux, make the AppImage executable with chmod +x Jan-*.AppImage if necessary.
  2. 02
    First launch
    Jan creates its data folder (~/jan on macOS/Linux, %APPDATA%\Jan on Windows). This is where the models, conversations, and config will live. Keep this in mind if you want to put it on a secondary SSD.
  3. 03
    Check Cortex
    At startup, Jan launches the Cortex service in the background. If a firewall window opens, allow the local connection (Cortex listens on 127.0.0.1, not on the network).
  4. 04
    Onboarding
    Jan suggests a starter model (typically a small Qwen or Granite). You can accept it to test in 2 minutes, or skip it to choose one yourself in the Hub.

#2. First model

Jan includes a model Hub that points to Hugging Face with prepared and tested GGUF versions. Open the “Hub” tab in the left sidebar to view the catalog.

To get started, three solid choices in French depending on your VRAM:

Small setup (8 GB RAM, no GPU)
Granite 4.2 3B Q4_K_M (≈2.2 GB). Very lightweight, enough for chatting, summarizing a short text, and brainstorming. Decent tokens/sec on CPU. Even lighter: Qwen 3.5 2B (≈1.9 GB).
Standard configuration (16 GB RAM, 6–8 GB GPU)
Qwen 3.5 9B Q4_K_M (≈6.6 GB). The 2026 8 GB choice: 256k context, vision, and an excellent FR/EN compromise for most everyday uses.
Comfortable setup (32 GB RAM, 12 GB+ GPU)
Gemma 4 12B (multimodal, ≈7.6 GB) or Mistral Small 24B in Q4_K_M (≈14 GB, good in French). Reasoning is clearly better than 8–9B models.

Click "Download" on the model page. Jan displays the progress and stores the GGUF file in its data folder. Once downloaded, the button becomes "Use"—one click loads the model into memory.

→
Import a GGUF you already have
If you already have local GGUF files (downloaded via huggingface-cli or obtained from a Ollama/LM Studio installation), you can import them into Jan via Settings → My Models → Import. Jan doesn’t duplicate the file; it references it.

#3. Getting Started

The Jan interface follows ChatGPT conventions: a conversation sidebar on the left, a chat area in the center, and a contextual settings panel on the right. Three things to know to be productive:

Threads
Each conversation is an independent “Thread” with its own assigned model. You can open multiple threads in parallel (useful for comparing the response of a 7B and a 14B to the same question).
Assistants
Assistants tab: create personas with a system prompt, temperature, and preassigned model. Handy for keeping an “Assistant code” (Qwen3-Coder, temperature 0.2) and an “Assistant writing” (Mistral Small, temperature 0.8) separate.
Generation parameters
Right panel—temperature, top-p, top-k, max tokens, penalty frequency. Adjustable per thread. The context window (n_ctx) is configured at the model level under Settings → My Models.
i
Conversations stored in plaintext
By default, Jan stores your threads in readable JSON files in the data folder. No encryption. If you process sensitive data, encrypt the disk (LUKS, FileVault, BitLocker) or place the Jan folder on an encrypted volume.

#4. Extensions and Cortex

Jan has an extension-based architecture. Its main components are themselves internal extensions (Inference Cortex Extension, Model Extension, etc.), which eventually makes it possible to swap them out. On the user side, the third-party extension ecosystem is still younger than VS Code's, but a few are worth checking out:

Alternative inference engines
Beyond the default Cortex/llama.cpp, you can enable extensions that point to a remote OpenAI-compatible endpoint (Ollama, vLLM, or even OpenAI cloud if you accept leaving the local setup).
Cortex standalone
Cortex ships with Jan but can run on its own. cortex run qwen3.5:2b in the CLI starts a server on the default port without the Jan frontend. Useful for scripting or deploying on a headless server.
Configurable Model Hub
You can add custom model sources (a private HF repo, an internal mirror) by editing the Hub settings. Useful in business environments for distributing in-house fine-tuned models.
Cortex in the CLI without Jan
# Lancer le serveur Cortex seul
cortex start

# Lister les modèles disponibles
cortex models list

# Charger et servir un modèle
cortex run qwen3.5:9b-gguf-q4-km

#5. API server mode

Jan's server mode exposes an OpenAI-compatible API on the local network. This transforms Jan from a simple chat client into a real backend for your scripts, n8n automations, or a VS Code copilot through Continue.dev.

  1. 01
    Enable the server
    Settings → Local API Server. Check "Start Local API Server." Jan starts Cortex in server mode on port 1337 by default (adjustable).
  2. 02
    Choose the served model
    All downloaded models are exposed. Each has an identifier you can use in the "model" field of requests (visible in the ID column of My Models).
  3. 03
    Quick curl test
    Verify that the server responds to a standard /v1/chat/completions call. The syntax is identical to OpenAI’s.
OpenAI-compatible API test
curl http://localhost:1337/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.5:9b-gguf-q4-km",
    "messages": [{"role": "user", "content": "Bonjour, qui es-tu ?"}],
    "temperature": 0.7
  }'
!
The server listens locally by default
Jan binds to 127.0.0.1, so it is inaccessible from the network. To share it with other machines (team, family), change the bind address to 0.0.0.0 in the settings—but then add authentication (at least a reverse proxy with basic auth); Jan has no built-in authentication.
Python client via the OpenAI SDK
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1337/v1",
    api_key="jan-local",  # n'importe quelle string, Jan ne vérifie pas
)

response = client.chat.completions.create(
    model="qwen3.5:9b-gguf-q4-km",
    messages=[
        {"role": "system", "content": "Tu réponds toujours en français."},
        {"role": "user", "content": "Explique-moi ce qu'est un LLM en deux phrases."},
    ],
)
print(response.choices[0].message.content)

#Jan vs LM Studio vs Msty

All three apps are desktop clients for local LLMs, with a clean UI and a local API server. The devil is in the details.

Jan
Open source (AGPL), strict offline-first, modular architecture with Cortex as a separate engine. Smaller model catalog than LM Studio. Ideal for those who want freedom and auditability.
LM Studio
Closed source but free. The richest Hugging Face catalog of the three (live search in HF). MCP support, native MLX on Mac, multi-token prediction. More features, but opaque.
Msty
Closed source, freemium (paid Aurum version). Best for RAG: document workspaces, split chat (compare two models side by side), and Knowledge Stacks. Less solid for pure chat, more oriented toward a “personal agent.”
→
How to choose
You want verifiable open source and simplicity: Jan. You want the maximum number of models and advanced settings: LM Studio. You want RAG over your documents without coding: Msty. None prevents you from using the others — they can even coexist (on different ports).

#Troubleshooting

The model won't load / OOM
Check the available VRAM/RAM. In Settings → My Models → [model] → Edit Parameters, lower n_gpu_layers to offload fewer layers to the GPU. Or switch to a smaller quantization (Q4_K_S instead of Q4_K_M, or IQ3_M).
Very low tokens/sec
Confirm that the GPU is actually being used: on NVIDIA, run nvidia-smi during generation; you should see Jan or cortex-server consuming VRAM. On Apple Silicon, Metal is enabled by default with no configuration required.
Cortex does not start (Windows)
Visual C++ Redistributable missing. Install vc_redist x64 from Microsoft. Also check that no other process is using port 1337 (change the port in settings if needed).
Linux: AppImage won't launch
Install libfuse2 (sudo apt install libfuse2 on Ubuntu 22.04+). For debugging, launch the AppImage from a terminal to see the errors.
Model download stuck at 99%
Stop and restart. Jan resumes the download. If the problem persists, download the GGUF manually from Hugging Face and import it via Settings → My Models → Import.
Answers in English despite a French prompt
Model not strong enough for consistent bilingual output. Add an explicit system prompt, « Always respond in French », to the assistant, or switch to a strong French model (Mistral Small 24B, Qwen 3.5, Gemma 4).

#Go further

Depending on where you want to go now:

Compare Jan with its direct competitors
“Ollama vs LM Studio vs Jan vs GPT4All” provides the detailed comparison for choosing the right tool for your profile.
Understanding Q4/Q5/Q8 quantizations
“Choosing your quantization (Q4, Q5, Q8, FP16)” explains the quality/memory tradeoffs—useful for making an informed choice in Jan’s Hub.
Run RAG on your documents
Jan does not have a robust native RAG solution. “Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” presents no-code alternatives that also connect to Jan’s API server.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.