Beginner 12 minInterfaces

LM Studio: complete 2026 tutorial (installation, API)

LM Studio is a desktop app that turns your machine into a local LLM server without touching the terminal. Version 0.3+ adds MCP support, a better GGUF engine, and a model-download experience directly from Hugging Face. Here’s a complete LM Studio 2026 tutorial: installation on all three OSes, your first model, an OpenAI-compatible server on port 1234, and an honest comparison with Ollama and Jan.

By Mohamed Meguedmi·Update 2026-09-12·Tested on Windows 11

#First result: download, loading, verification

The official documentation describes the Discover → download → Chat → model-loading workflow. Choose the file for your engine: GGUF for llama.cpp, or a compatible MLX distribution on Apple Silicon. A downloaded model is not loaded yet. Start with a short context, then ask a simple question before adding documents or an API server.

If memory is insufficient, reduce the context and choose a smaller file; control GPU offload. The model and context share the budget with the system on Apple Silicon. The Windows/macOS instructions are reviewed on documentation; the September 12 run was executed under Linux with Ollama, not with LM Studio.

Check installed models and estimate loading requirements
lms ls
# Remplacer MODEL_KEY par une clé réellement affichée :
lms load MODEL_KEY --context-length 4096 --estimate-only

#Why LM Studio in 2026

The Local AI Kit

LM Studio is installed and your first model responds. The rest is in the Local AI Kit: choose the model that fits your machine (ch. 3), take LM Studio further (ch. 5), then have it read your own documents (ch. 8).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

LM Studio fills a specific niche: users who want a real graphical app, not a CLI daemon. Open the app, find a model in the built-in search bar (which points to Hugging Face), click Download, and start chatting. Zero config files, zero command-line input.

Under the hood, LM Studio bundles llama.cpp with CUDA, Metal, Vulkan, and ROCm support, along with a dedicated MLX engine for Apple Silicon Macs. Version 0.3 introduced the “Power User” screen, exposing all parameters (KV cache type, flash attention, layer-by-layer GPU offload), and 0.3.x added native support for MCP servers—you can connect your local model to tools such as a filesystem or database without coding.

i
Licensing and commercial use
LM Studio is free for personal use. For commercial enterprise use, the "LM Studio at Work" form is now self-serve starting with 0.3.5—free, but required. No model telemetry is sent by default.

#Prerequisites

Windows
Windows 10/11 64-bit. AVX2 required (every CPU since 2014). CUDA support if NVIDIA GPU, Vulkan otherwise.
macOS
macOS 13.4+ on Apple Silicon (M1/M2/M3/M4). Intel Macs have no longer been supported since 0.3.
Linux
AppImage x86_64, glibc 2.35+ (Ubuntu 22.04, Fedora 36, recent Arch). ARM64 build available for SBCs.
RAM
16 GB is the comfortable minimum; 8 GB is limiting. For a 14B in Q4, plan for 12 GB free; for a 32B, 24 GB.
Disk
Plan for 30 to 50 GB: a 14B model in Q4 weighs ~9 GB, and you’ll download several.

#1. Installation on Windows, Mac, and Linux

Download the installer for your OS from the official website. No npm package, no Homebrew—it is a standard desktop app.

Official website
https://lmstudio.ai/download
  1. 01
    Windows
    Run LM-Studio-Setup.exe. SmartScreen may block the binary on first launch — click “More info,” then “Run anyway.” The app installs to %LOCALAPPDATA%\Programs\lm-studio.
  2. 02
    macOS
    Open the .dmg and drag LM Studio.app into /Applications. On first launch, Gatekeeper may ask you to allow the app in Settings → Privacy & Security. On Apple Silicon, the MLX engine is enabled by default.
  3. 03
    Linux
    Download the AppImage, make it executable with chmod +x LM_Studio-0.3.x.AppImage, then launch it. To integrate it into the menu, use AppImageLauncher or create a .desktop file manually.
Linux — first launch
chmod +x LM_Studio-0.3.x.AppImage
./LM_Studio-0.3.x.AppImage
→
2026 onboarding
On first launch, LM Studio 0.3+ offers a “User,” “Power User,” or “Developer” mode. Power User mode unlocks fine-grained controls (GPU layers, KV cache, flash attention); Developer mode adds the Local Server tab. Start with Power User—you can switch to Developer once the server API becomes useful.

#2. Download Qwen3-Coder in GGUF

The Discover tab (magnifying glass) is your interface to Hugging Face. Enter a model name, choose a quantization, and click Download. For this tutorial, we use Qwen3-Coder-30B-A3B, the reference open-weight coding model for users with 16 GB of VRAM or more.

  1. 01
    Search for the model
    In Discover, type "qwen3-coder". LM Studio displays compatible GGUF repositories (usually bartowski, unsloth, lmstudio-community). Prefer the lmstudio-community or bartowski repositories: their GGUFs are validated against the current engine version.
  2. 02
    Choose the quantization
    For 24 GB of VRAM, Q4_K_M is the sweet spot (about 99% of FP16 quality, about 17 GB in size). For 16 GB, drop to Q3_K_M. For 32 GB or a Mac Studio, use Q5_K_M or Q8_0 if you want maximum accuracy.
  3. 03
    Start the download
    Click Download. Progress appears in the Downloads tab. On a fiber connection, allow ~5 minutes for a 17 GB file.
  4. 04
    Load the model
    Switch to the Chat tab, open the selector at the top, and choose Qwen3-Coder. If you’re a Power User, adjust GPU Offload (by default LM Studio uses the max safe value), set Context Length (32k is comfortable for code), and enable Flash Attention.
→
VRAM by size (Q4_K_M)
Quick reference: 3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. If LM Studio refuses to load the model entirely onto the GPU, it switches to partial CPU offload—it works, but is 5 to 10× slower.

#3. First chat and GPU verification

With the model loaded, ask a question in Chat. LM Studio displays the tokens/sec, time to first token, and actual VRAM usage at the bottom of each response. It's more informative than running nvidia-smi in parallel.

Quick test
# Dans le chat LM Studio :
Écris une fonction Python qui parse un CSV streaming
sans charger tout le fichier en mémoire.

On a RTX 4090, Qwen3-Coder-30B-A3B in Q4 runs at around 90-120 tok/s thanks to its MoE architecture (3B active out of 30B). On an M4 Pro 48 GB, expect 35-50 tok/s with the MLX engine. On RTX 3060 12 GB, the 30B model does not fit—switch to Qwen 3.5 9B (~6.6 GB in Q4), which fits comfortably and leaves room for the context.

i
“Just Show Me The Answer” mode
By default, reasoning models display their chain of thought in a collapsible section. In a chat's Settings, you can disable thinking display to see only the final answer. Note: Qwen 3.8 27B tends to overthink with its default reasoning setting—switch it to “low” effort if responses are taking unnecessarily long.

#4. OpenAI-compatible server on port 1234

This is the feature many people miss: LM Studio exposes an HTTP server compatible with the OpenAI API, by default on http://localhost:1234. You can therefore connect any OpenAI client (Python SDK, LangChain, Continue.dev, Cursor, your own app) to LM Studio without changing a line of business logic—just the base URL and key.

  1. 01
    Enable Developer Mode
    At the bottom left, change the user role to "Developer". The Local Server tab (terminal icon) appears in the left sidebar.
  2. 02
    Start the server
    Local Server tab → choose the model to expose from the top menu → click Start Server. By default, the server listens on localhost:1234. Check "Serve on Local Network" if you want to expose it to other machines on your LAN.
  3. 03
    Check the endpoint
    The server exposes the /v1/models, /v1/chat/completions, /v1/completions, /v1/embeddings routes (if an embedding model is loaded). First, test the model list.
API test
curl http://localhost:1234/v1/models

curl http://localhost:1234/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3-coder-30b-a3b",
    "messages": [{"role": "user", "content": "Bonjour"}],
    "stream": false
  }'
Official Python client
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="lm-studio",  # ignorée, n'importe quelle chaîne
)

resp = client.chat.completions.create(
    model="qwen3-coder-30b-a3b",
    messages=[{"role": "user", "content": "Refactore ce code en idiomatic Rust : ..."}],
)
print(resp.choices[0].message.content)
→
JIT model loading
Enable “Just-in-Time Model Loading” in the server settings. LM Studio then automatically loads the model requested by the API request — useful if you switch between multiple models from different clients.

#5. Native MCP support

Since 0.3.17, LM Studio has supported the Model Context Protocol (MCP)—the standard introduced by Anthropic in 2024 for connecting external tools to an LLM. In practice, your local model can access your file system, execute SQL, and call an API, provided an MCP server is configured.

Configuration is done in the mcp.json file, accessible from the Program → Edit mcp.json menu. Here is an example that adds Anthropic's official filesystem server and a SQLite server.

mcp.json
{
  "mcpServers": {
    "filesystem": {
      "command": "npx",
      "args": [
        "-y",
        "@modelcontextprotocol/server-filesystem",
        "/home/user/Documents"
      ]
    },
    "sqlite": {
      "command": "uvx",
      "args": ["mcp-server-sqlite", "--db-path", "/home/user/data.db"]
    }
  }
}

Once registered, the MCP tools appear under the plug icon in the chat. You can enable them as needed. The model then decides when to call them—useful for creating a mini local agent without coding, such as “summarize all the PDFs in the Research folder” or “how many rows are in the users table?”

!
Which model for tool use
Not all models are good at tool calling. Qwen3-Coder 30B-A3B, GLM 4.7 Flash (excellent at agents), Devstral 24B, and Mistral Small 24B work well, as do any models explicitly trained for function calling. Small models (< 7B) often hallucinate tool names—test before deploying.

#LM Studio vs. Ollama vs. Jan

The three tools target the same goal (running an LLM locally without the cloud) with different philosophies.

LM Studio
Feature-rich desktop app, integrated HF search, API server on :1234, MCP support, and MLX on Mac. Closed source but free. Ideal if you want everything in one place with a serious GUI.
Ollama
Lightweight daemon + CLI, listening on :11434, huge library of preconfigured models (ollama run qwen3.5:9b), stable API, integrations everywhere (Cline, Open WebUI, n8n). Open source. Ideal if you're comfortable in the terminal and want a headless setup.
Jan
Open-source desktop app (AGPL), with a radically "offline-first" philosophy, newer and less feature-rich than LM Studio. Nice if open source matters to you and you want a desktop app.
→
Winning combination
Many advanced users run Ollama as a server daemon (24/7, without a GUI) and use LM Studio as a client/explorer to test new models before pushing them into Ollama. The two can coexist—they don't listen on the same port.

#Troubleshooting

The model won't load (OOM)
Manually reduce GPU Offload in the chat Settings. If offloading 100% to the GPU causes a crash, lower it to 80%—a few layers will use RAM instead, which is slower but works.
Very low tokens/sec
Make sure Flash Attention is enabled and KV cache type is set to Q8_0 (not FP16). In Power User, these settings are in the chat's right sidebar.
The API server is not responding
Confirm that you are in Developer Mode and that the Start Server button is green. On Windows, the firewall may block port 1234 — allow LM Studio through Windows Defender Firewall.
Model not found in search
LM Studio lists only GGUF repositories compatible with its version of llama.cpp. For very recent models, update LM Studio (Help → Check for Updates) or download the GGUF manually and place it in the models folder.
Linux: AppImage refuses to launch
Install libfuse2 (sudo apt install libfuse2 on Ubuntu 22.04+). On Wayland, export QT_QPA_PLATFORM=xcb if the app does not display correctly.

#Go further

Three natural directions depending on what you want to do next:

Compare in more detail
The “LM Studio vs Ollama: which should you choose in 2026?” guide compares them feature by feature, with a decision matrix based on your profile.
Connect LM Studio to your PDFs
“Local RAG with LM Studio: chat with your documents” explains the Chat with Documents feature, its limitations, and when to switch to AnythingLLM.
Accelerating inference even further
“LM Studio and MTP (Multi-Token Prediction)” explains how to enable MTP on compatible models to gain 30 to 80% more throughput.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.