LM Studio for beginners: first local chat in 10 minutes
LM Studio is the closest equivalent to ChatGPT for local AI: a graphical application, an integrated store for finding and downloading models, and a chat interface that resembles what you already know. Zero command-line work. Perfect for getting started without a terminal.
#What is LM Studio?
A desktop application (Windows, macOS, Linux) that packages four things into a single interface: a model browser connected to Hugging Face, a download manager, an inference engine (llama.cpp under the hood), and a chat.
The tool is free for personal use. For commercial enterprise use, a license has been required since version 0.3.
#1. Installation
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- 01Choose your platformWindows .exe, macOS .dmg (Apple Silicon or Intel), Linux .AppImage. The binaries are between 400 and 600 MB—normal, since llama.cpp and its dependencies are included.
- 02Run the installerOn Windows, installation takes place in %LOCALAPPDATA% (no admin rights required). On Mac, drag it to Applications. On Linux, make the AppImage executable: chmod +x LM_Studio_*.AppImage.
- 03First launchLM Studio asks which interface level to display (Power user or Developer). Power user is more than enough to get started—you can switch later.
#2. Tour of the premises
The left sidebar contains the 5 essential views:
- 💬 Chat
- The main conversation interface. Looks like ChatGPT.
- 🔍 Discover
- The model “store.” Live search on Hugging Face, one-click downloads.
- 📁 My Models
- All the models you’ve downloaded, along with their sizes and quantizations.
- 💻 Developer
- The local OpenAI-compatible server mode (see below).
- ⚙️ Settings
- Global settings: model folder, theme, integrations.
#3. Download your first model
Click 🔍 Discover. The search bar searches directly on Hugging Face. For a first try in 2026, enter Qwen 3.5 9B—the best choice for 8 GB of VRAM (256k context, vision, Apache 2.0 license).
- 01Check the compatibility tagsTo the right of each result, LM Studio indicates Full GPU Offload possible, Partial GPU Offload, or Likely too large. It's based on your detected hardware—you only need to choose a green one.
- 02Choose the right quantizationFor a 9B model such as Qwen 3.5 9B on 8 GB of VRAM: Q4_K_M is the sweet spot. For 12 GB+: Q5_K_M or Q6_K (or even the 9B’s Q8_0). For 16 GB+: Q8_0 for maximum quality.
- 03Click DownloadBetween 4 and 8 GB depending on the quantization. LM Studio displays progress live.
#4. First conversation
- 01Return to the Chat tabAt the top, click the "Select a model to load" selector and choose the model you just downloaded.
- 02Adjust GPU offloadLM Studio provides a slider: 0 = CPU only (slow but always works), max = GPU only (fast but crashes if VRAM is insufficient). Set the slider to the maximum value proposed by default.
- 03Click Load modelLoading takes 5 to 30 seconds depending on the model size and your disk speed.
- 04Enter your questionIn the field at the bottom. Press Enter to send. The model responds in streaming mode, one token at a time.
#5. Tune response quality
Right panel during a conversation. Three settings are enough to change 90% of the behavior:
- Temperature
- 0.2 = factual, deterministic. 0.7 = balanced (default). 1.2 = creative, sometimes erratic. Increase it if the model is too repetitive.
- Context Length
- How many tokens the model “sees” backward. 4096 by default. Increase it to 8192 or 16384 for long documents—be careful, this consumes VRAM.
- System Prompt
- The invisible message that defines the role and rules. E.g.: "You are a French legal assistant. Always answer by citing the applicable law article."
#6. Local server mode (API)
LM Studio exposes an HTTP server compatible with the OpenAI API. Any tool that knows how to communicate with ChatGPT can therefore run locally.
- 01Developer tabSelect a model and click Start Server (on port 1234 by default).
- 02Test with curlFrom a terminal, a simple GET verifies that the server responds.
In your Python scripts, simply change base_url to http://localhost:1234/v1 and use any dummy API key. Anything that works with OpenAI works.
#Limitations to know
- Not fully open source
- The app itself is proprietary; only the llama.cpp engine behind it is open source. For a 100% free and open-source option, prefer Jan or Ollama.
- Automatic updates
- LM Studio updates itself. Convenient, but unexpected in a controlled production environment.
- Paid commercial license
- For business use, read the terms. Personal and educational use remains free.
- No native scripting
- Unlike Ollama, it has no CLI. Automations go exclusively through the API server.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.