Beginner 9 minInterfaces

LM Studio on Linux: Ubuntu, Debian, Arch, Fedora (2026)

LM Studio on Linux is an ~600 MB AppImage that bundles a graphical interface, a model store connected to Hugging Face, a llama.cpp inference engine, and an OpenAI-compatible server. No system installation, no daemon, no repository to add — one executable file, and that’s it. This guide covers installation, downloading a GGUF, starting the local server, and Linux-specific pitfalls (permissions, FUSE, GPU).

By Mohamed Meguedmi·Update 2026-08-27·Tested on Ubuntu 24.04
i
In brief
LM Studio on Linux is an AppImage of about 600 MB: no system installation, no daemon, just one executable file. · Three steps are enough: download the file, make it executable (chmod +x), then run it. · Requirements: a recent distribution with glibc 2.35+, and the libfuse2 package on Ubuntu 22.04 and later. · The AppImage automatically detects the GPU (CUDA, ROCm, or Vulkan depending on your card) to accelerate inference.

#Why LM Studio on Linux

On Linux, the usual choice for a local LLM is Ollama: a systemd daemon, a clean CLI, listening by default on localhost:11434. It's excellent for the server. LM Studio takes a different angle — the graphical workstation.

A complete interface without a terminal
Chat, Hugging Face model browser, download manager, sampling settings: everything in a GUI. Useful if you want to quickly test several models before moving to production.
A one-click OpenAI-compatible server
The Developer tab exposes a http://localhost:1234/v1 endpoint that speaks the OpenAI protocol. Any SDK (openai-python, LangChain, Continue.dev) can call it without changing a single line.
Zero system installation
The AppImage runs in user space. No sudo, no package to install, no third-party repository. For machines locked down by IT or unusual distributions, that’s valuable.
Native GGUF format
LM Studio reads only GGUF (the llama.cpp format). It's the same format used under the hood by Ollama, so 100% of popular models are available.
i
LM Studio and Ollama coexist very well
There's nothing stopping you from having both. Ollama as a daemon for your scripts and integrations, LM Studio for interactive chat and discovering new models. They don't use the same ports (11434 vs 1234) and don't interfere with each other.

#Prerequisites

The Local AI Kit

LM Studio runs on your Linux system. The Local AI Kit takes over: advanced settings for LM Studio (ch. 5) and, when something goes wrong, a symptom-based troubleshooting tree—slowness, graphics card ignored, model forgetting everything (ch. 14).

  • Lifetime online access
  • PDF + files
  • Lifetime updates
Recent Linux distribution
Ubuntu 22.04+, Fedora 38+, Debian 12+, Arch, or derivatives. LM Studio is tested on the main distros, but the AppImage is portable and works elsewhere (openSUSE, Linux Mint, Pop!_OS).
glibc 2.35+
The AppImage bundles its graphics dependencies but not glibc. If you're on a very old distribution (CentOS 7, Debian 10), it won't run.
FUSE
AppImages use FUSE to mount themselves read-only. Most distros install it by default. On Ubuntu 22.04+, you will need to explicitly install the libfuse2 package.
8 GB of RAM minimum
To run a 7B model in Q4. 16 GB for comfortable use, 32 GB+ for 14B models or to keep your IDE open alongside it.
GPU (optional but recommended)
NVIDIA with recent drivers (CUDA 12+), AMD with ROCm 6.x, or Intel Arc via Vulkan. Without a GPU, it runs on the CPU — functional for 3B models, slow for 7B models.
!
Ubuntu 22.04 and FUSE
Ubuntu 22.04 dropped libfuse2 by default, which breaks many AppImages with an obscure message such as "dlopen(): error loading libfuse.so.2". Install the package with sudo apt install libfuse2 before launching LM Studio.

#1. Retrieve the AppImage

The official site automatically detects your OS and offers the right version. Make sure you're downloading from lmstudio.ai—there are questionable mirrors.

Official website
https://lmstudio.ai
  1. 01
    Download the .AppImage file
    The binary is between 500 and 700 MB. It includes llama.cpp, the Electron UI, and the entire runtime. Conventionally, place it in ~/Applications/ or ~/.local/bin/—in reality, anywhere is fine; the AppImage is self-contained.
  2. 02
    Make it executable
    With the chmod +x command. Without it, the file is just an inert archive.
  3. 03
    Lancez-le
    Double-click from your file manager, or run it from the command line. On first launch, LM Studio extracts its contents to ~/.cache/AppImage/ — this is normal, and it only happens again after each update.
Quick CLI installation
mkdir -p ~/Applications
cd ~/Applications
# Remplacez le nom par celui du fichier téléchargé
chmod +x LM_Studio-*.AppImage
./LM_Studio-*.AppImage
→
Add LM Studio to the application menu
Install appimaged or AppImageLauncher: these tools automatically detect AppImages in ~/Applications/ and create .desktop entries so they appear in your menu like a native app. AppImageLauncher even handles updates.

#2. First launch

On first launch, LM Studio asks you to choose an interface level: User (chat only), Power User (chat + settings), Developer (everything, including the server). Power User is enough to explore; you can switch to Developer for the local server later from Settings → UI Mode.

The left sidebar organizes the views:

Chat
The main interface, such as ChatGPT. This is what you will use 90% of the time.
Discover
The model store. Live search on Hugging Face, with a hardware compatibility indicator (Full GPU Offload / Partial / Likely too large) based on your detected VRAM.
My Models
All your downloaded models, with their sizes and quantizations. Useful for cleaning things up.
Developer
The local OpenAI-compatible server. Visible only in Power User or Developer mode.

#3. Download a GGUF model

LM Studio only reads the GGUF format (llama.cpp's binary format, successor to GGML). If a model exists on Hugging Face but not in GGUF, you won't be able to load it directly. The vast majority of popular models (Qwen 3.5, Gemma 4, Mistral, Granite 4.2, DeepSeek, gpt-oss) have community-maintained GGUF versions.

  1. 01
    Open Discover and search for a model
    For a first try, type Qwen 3.5 9B or Granite 4.2 8B—the safe 2026 choices in the 6-8 GB VRAM range. LM Studio lists the available GGUF variants, generally published by bartowski, unsloth, or lmstudio-community.
  2. 02
    Read the compatibility tags
    To the right of each variant, a green/orange/red indicator: Full GPU Offload means the entire model fits in your VRAM. Partial = some of it will go into RAM (slower). Likely too large = forget it.
  3. 03
    Choose the quantization
    Q4_K_M is the sweet spot for most use cases (small quality loss, ~50% of FP16 size). Q5_K_M if you have 30–40% more VRAM. Q8_0 for maximum quality if you have plenty of space. FP16 only for fine-tuning or comparison.
  4. 04
    Click Download
    Between 4 and 8 GB for a 7B, depending on the quantization. The download runs in the background, so you can start several in parallel.
→
Where models are stored
By default, LM Studio stores models in ~/.cache/lm-studio/models/. On Linux, this directory can grow quickly (50-200 GB with a few models). If your /home is on a small SSD, move the directory via Settings → Model Directory to a larger drive.

Once the model has downloaded, return to Chat, select it from the dropdown menu at the top, move the GPU offload slider to the maximum available setting, and click Load model. Loading takes 5 to 30 seconds depending on the size and speed of the disk.

#4. Launch the local server on port 1234

This is where LM Studio goes beyond simple graphical chat. The Developer tab exposes an HTTP server that speaks exactly the same protocol as the OpenAI API. Any OpenAI client (openai-python, LangChain, LlamaIndex, Continue.dev) can connect to it without modification—just change the base URL.

  1. 01
    Switch to Developer mode
    Settings → UI Mode → Developer. The Developer tab (terminal icon) appears in the sidebar.
  2. 02
    Select a model
    Menu at the top of the Developer tab. The model must already be downloaded. You can load several in parallel if VRAM permits.
  3. 03
    Adjust Context Length and GPU Offload
    Set GPU offload as high as VRAM allows. The default context length of 4096 is sufficient for most uses; increase it to 8192 or 16384 for RAG or long documents.
  4. 04
    Click Start Server
    The server listens on 127.0.0.1:1234 by default. You can configure the port right next to it if 1234 is already in use.
Test the server
# Liste des modèles chargés
curl http://localhost:1234/v1/models

# Une complétion chat minimale
curl http://localhost:1234/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "local-model",
    "messages": [
      {"role":"user","content":"Capitale de l'\''Italie ?"}
    ],
    "temperature": 0.2
  }'

On the Python side, an official openai client is enough. The API key can be any string—LM Studio does not verify it.

Python client
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="lm-studio",  # placeholder, non vérifié
)

resp = client.chat.completions.create(
    model="local-model",
    messages=[
        {"role": "system", "content": "Tu réponds en une phrase."},
        {"role": "user",   "content": "Qu'est-ce qu'un GGUF ?"},
    ],
    stream=True,
)

for chunk in resp:
    print(chunk.choices[0].delta.content or "", end="", flush=True)
i
Live logs
The Developer tab displays the Server Logs in real time: every incoming request, generation time, and token count. Essential for debugging an integration.

To expose the server on the LAN (other team workstations), change Host from 127.0.0.1 to 0.0.0.0 in the server settings. Warning: LM Studio has no built-in authentication. On a shared network, put Caddy or Nginx with basic auth in front, or restrict access with the firewall.

#5. GPU acceleration on Linux

This is the trickiest part of LM Studio on Linux. The AppImage bundles several backends (CPU, CUDA, ROCm, Vulkan) and chooses automatically at launch—but autodetection sometimes gets it wrong depending on the installed drivers.

NVIDIA (CUDA)
Proprietary 535+ drivers installed (nvidia-smi must work). CUDA 12.x is bundled in the AppImage, so you do not need to install it separately. Check under Settings → Hardware that the GPU is listed and the backend is set to CUDA.
AMD (ROCm)
ROCm 6.x drivers installed on the host (Radeon RX 6800 XT and newer are officially supported; RX 6700 XT and RX 7600 often work with HSA_OVERRIDE_GFX_VERSION). Your user must belong to the render and video groups.
Intel Arc / iGPU
Vulkan backend, decent performance on Arc A770/A750 but worse than NVIDIA/AMD. Install vulkan-tools to verify with vulkaninfo that the GPU is detected.
No GPU
Automatic CPU fallback. For a Granite 4.2 8B Q4 on a Ryzen 7, expect 5-8 tokens/sec. A 3B model such as Qwen 3.5 2B or Granite 4.2 3B remains very usable. Above 14B, patience becomes an issue.

The GPU offload slider in Chat or Developer controls how many model layers go to the GPU. Set it as high as VRAM allows. Check actual usage in a terminal:

Monitor VRAM
# NVIDIA
nvidia-smi -l 1

# AMD (radeontop ou rocm-smi)
rocm-smi --showmeminfo vram
VRAM guidelines by model size (Q4)
3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. These figures increase slightly if you raise the context length beyond 4096.
Reference GPUs
RTX 3060 12GB comfortably covers 7B-14B. RTX 4070 12GB: same, with headroom. RTX 4080 16GB: 14B in Q5, the start of 24B. RTX 4090 24GB: 32B in Q4, 70B in Q3. Mac M4 Pro with 24-48GB unified memory: the same thing on the Apple Silicon side, but through the macOS version of LM Studio.
!
The silent CPU fallback trap
If there isn't enough VRAM for the model at the requested context length, LM Studio moves some layers to RAM without always making this clear. You get 2–5 tokens/sec instead of 40+. Check Settings → Hardware or nvidia-smi during generation—if VRAM isn't saturated, part of the workload is running on the CPU.

#Common pitfalls

The AppImage refuses to start
Run it from a terminal to see the error message. Most common: libfuse2 missing on Ubuntu 22.04+ (sudo apt install libfuse2 fixes the problem). Next: glibc is too old (LM Studio requires 2.35+, so no CentOS 7 or Debian 10).
The GPU is not being used
Check nvidia-smi (or rocm-smi). If the GPU isn’t listed: drivers are missing or too old. If it is listed but LM Studio remains on the CPU: Settings → Hardware lets you force the backend (CUDA/ROCm/Vulkan). On AMD, the user must be in the render and video groups—log out and back in after running usermod.
The model is slow despite the GPU
VRAM is saturated and partial offloading is happening silently. Reduce the context length, lower the GPU offload slider to measure the threshold, or use a more aggressive quantization (Q5 → Q4 → Q3).
Port 1234 is already in use
Often occupied by another development service. Change the port in Developer → Settings, or kill the process: ss -tulpn | grep 1234. There is no need to touch the firewall for localhost use.
Breaking automatic updates
LM Studio updates itself by default. In a controlled environment (enterprise, CI/CD integration), disable this via Settings → Updates. Note the version number that works so you can roll back if needed.
No CLI autocompletion
The AppImage does not provide a rich CLI. For serious scripting, the server’s HTTP API (port 1234) remains the cleanest option. If you need a real CLI, llama.cpp or Ollama are better suited.

#Go further

You have LM Studio installed, a GGUF loaded, and an OpenAI-compatible server running. The natural next directions are:

Deep dive into the API server
The Transformer LM Studio as an API server guide covers streaming, multiple models in parallel, secure LAN exposure, and performance tuning.
Compare with Ollama on Linux
The guide Installing Ollama on Linux covers the other popular stack, which is more daemon/server-oriented. For many users, the combination of Ollama (daemon) + LM Studio (client interface on top) is ideal.
Choosing the right quantization
The Choosing Your Quantization (Q4, Q5, Q8, FP16) guide visually compares actual quality loss and helps weigh quality against VRAM, especially when downloading your own GGUFs.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.