Beginner 5 minLM Studio

LM Studio for beginners: first local chat in 10 minutes

LM Studio is the closest equivalent to ChatGPT for local AI: a graphical application, an integrated store for finding and downloading models, and a chat interface that resembles what you already know. Zero command-line work. Perfect for getting started without a terminal.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#What is LM Studio?

A desktop application (Windows, macOS, Linux) that packages four things into a single interface: a model browser connected to Hugging Face, a download manager, an inference engine (llama.cpp under the hood), and a chat.

The tool is free for personal use. For commercial enterprise use, a license has been required since version 0.3.

i
LM Studio vs Ollama
Ollama is CLI-first, optimized for scripts and APIs. LM Studio is GUI-first, optimized for exploration. Both are legitimate—many advanced users have both installed.

#1. Installation

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Official website
https://lmstudio.ai
  1. 01
    Choose your platform
    Windows .exe, macOS .dmg (Apple Silicon or Intel), Linux .AppImage. The binaries are between 400 and 600 MB—normal, since llama.cpp and its dependencies are included.
  2. 02
    Run the installer
    On Windows, installation takes place in %LOCALAPPDATA% (no admin rights required). On Mac, drag it to Applications. On Linux, make the AppImage executable: chmod +x LM_Studio_*.AppImage.
  3. 03
    First launch
    LM Studio asks which interface level to display (Power user or Developer). Power user is more than enough to get started—you can switch later.

#2. Tour of the premises

The left sidebar contains the 5 essential views:

💬 Chat
The main conversation interface. Looks like ChatGPT.
🔍 Discover
The model “store.” Live search on Hugging Face, one-click downloads.
📁 My Models
All the models you’ve downloaded, along with their sizes and quantizations.
💻 Developer
The local OpenAI-compatible server mode (see below).
⚙️ Settings
Global settings: model folder, theme, integrations.

#3. Download your first model

Click 🔍 Discover. The search bar searches directly on Hugging Face. For a first try in 2026, enter Qwen 3.5 9B—the best choice for 8 GB of VRAM (256k context, vision, Apache 2.0 license).

  1. 01
    Check the compatibility tags
    To the right of each result, LM Studio indicates Full GPU Offload possible, Partial GPU Offload, or Likely too large. It's based on your detected hardware—you only need to choose a green one.
  2. 02
    Choose the right quantization
    For a 9B model such as Qwen 3.5 9B on 8 GB of VRAM: Q4_K_M is the sweet spot. For 12 GB+: Q5_K_M or Q6_K (or even the 9B’s Q8_0). For 16 GB+: Q8_0 for maximum quality.
  3. 03
    Click Download
    Between 4 and 8 GB depending on the quantization. LM Studio displays progress live.
→
The "GGUF" tag
LM Studio can only read the GGUF format (the successor to GGML). If a model appears without a GGUF variant, you need to convert it yourself or avoid it.

#4. First conversation

  1. 01
    Return to the Chat tab
    At the top, click the "Select a model to load" selector and choose the model you just downloaded.
  2. 02
    Adjust GPU offload
    LM Studio provides a slider: 0 = CPU only (slow but always works), max = GPU only (fast but crashes if VRAM is insufficient). Set the slider to the maximum value proposed by default.
  3. 03
    Click Load model
    Loading takes 5 to 30 seconds depending on the model size and your disk speed.
  4. 04
    Enter your question
    In the field at the bottom. Press Enter to send. The model responds in streaming mode, one token at a time.

#5. Tune response quality

Right panel during a conversation. Three settings are enough to change 90% of the behavior:

Temperature
0.2 = factual, deterministic. 0.7 = balanced (default). 1.2 = creative, sometimes erratic. Increase it if the model is too repetitive.
Context Length
How many tokens the model “sees” backward. 4096 by default. Increase it to 8192 or 16384 for long documents—be careful, this consumes VRAM.
System Prompt
The invisible message that defines the role and rules. E.g.: "You are a French legal assistant. Always answer by citing the applicable law article."
!
Misleading presets
LM Studio includes several presets (Creative, Balanced, Precise). They change several settings at once. Convenient, but read what each one changes before applying them blindly.

#6. Local server mode (API)

LM Studio exposes an HTTP server compatible with the OpenAI API. Any tool that knows how to communicate with ChatGPT can therefore run locally.

  1. 01
    Developer tab
    Select a model and click Start Server (on port 1234 by default).
  2. 02
    Test with curl
    From a terminal, a simple GET verifies that the server responds.
Verification
curl http://localhost:1234/v1/models
Minimal chat
curl http://localhost:1234/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "local-model",
    "messages": [{"role":"user","content":"Bonjour"}]
  }'

In your Python scripts, simply change base_url to http://localhost:1234/v1 and use any dummy API key. Anything that works with OpenAI works.

#Limitations to know

Not fully open source
The app itself is proprietary; only the llama.cpp engine behind it is open source. For a 100% free and open-source option, prefer Jan or Ollama.
Automatic updates
LM Studio updates itself. Convenient, but unexpected in a controlled production environment.
Paid commercial license
For business use, read the terms. Personal and educational use remains free.
No native scripting
Unlike Ollama, it has no CLI. Automations go exclusively through the API server.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.