LM Studio on Linux: Ubuntu, Debian, Arch, Fedora (2026)
LM Studio on Linux is an ~600 MB AppImage that bundles a graphical interface, a model store connected to Hugging Face, a llama.cpp inference engine, and an OpenAI-compatible server. No system installation, no daemon, no repository to add — one executable file, and that’s it. This guide covers installation, downloading a GGUF, starting the local server, and Linux-specific pitfalls (permissions, FUSE, GPU).
#Why LM Studio on Linux
On Linux, the usual choice for a local LLM is Ollama: a systemd daemon, a clean CLI, listening by default on localhost:11434. It's excellent for the server. LM Studio takes a different angle — the graphical workstation.
- A complete interface without a terminal
- Chat, Hugging Face model browser, download manager, sampling settings: everything in a GUI. Useful if you want to quickly test several models before moving to production.
- A one-click OpenAI-compatible server
- The Developer tab exposes a http://localhost:1234/v1 endpoint that speaks the OpenAI protocol. Any SDK (openai-python, LangChain, Continue.dev) can call it without changing a single line.
- Zero system installation
- The AppImage runs in user space. No sudo, no package to install, no third-party repository. For machines locked down by IT or unusual distributions, that’s valuable.
- Native GGUF format
- LM Studio reads only GGUF (the llama.cpp format). It's the same format used under the hood by Ollama, so 100% of popular models are available.
#Prerequisites
LM Studio runs on your Linux system. The Local AI Kit takes over: advanced settings for LM Studio (ch. 5) and, when something goes wrong, a symptom-based troubleshooting tree—slowness, graphics card ignored, model forgetting everything (ch. 14).
- Lifetime online access
- PDF + files
- Lifetime updates
- Recent Linux distribution
- Ubuntu 22.04+, Fedora 38+, Debian 12+, Arch, or derivatives. LM Studio is tested on the main distros, but the AppImage is portable and works elsewhere (openSUSE, Linux Mint, Pop!_OS).
- glibc 2.35+
- The AppImage bundles its graphics dependencies but not glibc. If you're on a very old distribution (CentOS 7, Debian 10), it won't run.
- FUSE
- AppImages use FUSE to mount themselves read-only. Most distros install it by default. On Ubuntu 22.04+, you will need to explicitly install the libfuse2 package.
- 8 GB of RAM minimum
- To run a 7B model in Q4. 16 GB for comfortable use, 32 GB+ for 14B models or to keep your IDE open alongside it.
- GPU (optional but recommended)
- NVIDIA with recent drivers (CUDA 12+), AMD with ROCm 6.x, or Intel Arc via Vulkan. Without a GPU, it runs on the CPU — functional for 3B models, slow for 7B models.
#1. Retrieve the AppImage
The official site automatically detects your OS and offers the right version. Make sure you're downloading from lmstudio.ai—there are questionable mirrors.
- 01Download the .AppImage fileThe binary is between 500 and 700 MB. It includes llama.cpp, the Electron UI, and the entire runtime. Conventionally, place it in ~/Applications/ or ~/.local/bin/—in reality, anywhere is fine; the AppImage is self-contained.
- 02Make it executableWith the chmod +x command. Without it, the file is just an inert archive.
- 03Lancez-leDouble-click from your file manager, or run it from the command line. On first launch, LM Studio extracts its contents to ~/.cache/AppImage/ — this is normal, and it only happens again after each update.
#2. First launch
On first launch, LM Studio asks you to choose an interface level: User (chat only), Power User (chat + settings), Developer (everything, including the server). Power User is enough to explore; you can switch to Developer for the local server later from Settings → UI Mode.
The left sidebar organizes the views:
- Chat
- The main interface, such as ChatGPT. This is what you will use 90% of the time.
- Discover
- The model store. Live search on Hugging Face, with a hardware compatibility indicator (Full GPU Offload / Partial / Likely too large) based on your detected VRAM.
- My Models
- All your downloaded models, with their sizes and quantizations. Useful for cleaning things up.
- Developer
- The local OpenAI-compatible server. Visible only in Power User or Developer mode.
#3. Download a GGUF model
LM Studio only reads the GGUF format (llama.cpp's binary format, successor to GGML). If a model exists on Hugging Face but not in GGUF, you won't be able to load it directly. The vast majority of popular models (Qwen 3.5, Gemma 4, Mistral, Granite 4.2, DeepSeek, gpt-oss) have community-maintained GGUF versions.
- 01Open Discover and search for a modelFor a first try, type Qwen 3.5 9B or Granite 4.2 8B—the safe 2026 choices in the 6-8 GB VRAM range. LM Studio lists the available GGUF variants, generally published by bartowski, unsloth, or lmstudio-community.
- 02Read the compatibility tagsTo the right of each variant, a green/orange/red indicator: Full GPU Offload means the entire model fits in your VRAM. Partial = some of it will go into RAM (slower). Likely too large = forget it.
- 03Choose the quantizationQ4_K_M is the sweet spot for most use cases (small quality loss, ~50% of FP16 size). Q5_K_M if you have 30–40% more VRAM. Q8_0 for maximum quality if you have plenty of space. FP16 only for fine-tuning or comparison.
- 04Click DownloadBetween 4 and 8 GB for a 7B, depending on the quantization. The download runs in the background, so you can start several in parallel.
Once the model has downloaded, return to Chat, select it from the dropdown menu at the top, move the GPU offload slider to the maximum available setting, and click Load model. Loading takes 5 to 30 seconds depending on the size and speed of the disk.
#4. Launch the local server on port 1234
This is where LM Studio goes beyond simple graphical chat. The Developer tab exposes an HTTP server that speaks exactly the same protocol as the OpenAI API. Any OpenAI client (openai-python, LangChain, LlamaIndex, Continue.dev) can connect to it without modification—just change the base URL.
- 01Switch to Developer modeSettings → UI Mode → Developer. The Developer tab (terminal icon) appears in the sidebar.
- 02Select a modelMenu at the top of the Developer tab. The model must already be downloaded. You can load several in parallel if VRAM permits.
- 03Adjust Context Length and GPU OffloadSet GPU offload as high as VRAM allows. The default context length of 4096 is sufficient for most uses; increase it to 8192 or 16384 for RAG or long documents.
- 04Click Start ServerThe server listens on 127.0.0.1:1234 by default. You can configure the port right next to it if 1234 is already in use.
On the Python side, an official openai client is enough. The API key can be any string—LM Studio does not verify it.
To expose the server on the LAN (other team workstations), change Host from 127.0.0.1 to 0.0.0.0 in the server settings. Warning: LM Studio has no built-in authentication. On a shared network, put Caddy or Nginx with basic auth in front, or restrict access with the firewall.
#5. GPU acceleration on Linux
This is the trickiest part of LM Studio on Linux. The AppImage bundles several backends (CPU, CUDA, ROCm, Vulkan) and chooses automatically at launch—but autodetection sometimes gets it wrong depending on the installed drivers.
- NVIDIA (CUDA)
- Proprietary 535+ drivers installed (nvidia-smi must work). CUDA 12.x is bundled in the AppImage, so you do not need to install it separately. Check under Settings → Hardware that the GPU is listed and the backend is set to CUDA.
- AMD (ROCm)
- ROCm 6.x drivers installed on the host (Radeon RX 6800 XT and newer are officially supported; RX 6700 XT and RX 7600 often work with HSA_OVERRIDE_GFX_VERSION). Your user must belong to the render and video groups.
- Intel Arc / iGPU
- Vulkan backend, decent performance on Arc A770/A750 but worse than NVIDIA/AMD. Install vulkan-tools to verify with vulkaninfo that the GPU is detected.
- No GPU
- Automatic CPU fallback. For a Granite 4.2 8B Q4 on a Ryzen 7, expect 5-8 tokens/sec. A 3B model such as Qwen 3.5 2B or Granite 4.2 3B remains very usable. Above 14B, patience becomes an issue.
The GPU offload slider in Chat or Developer controls how many model layers go to the GPU. Set it as high as VRAM allows. Check actual usage in a terminal:
- VRAM guidelines by model size (Q4)
- 3B ≈ 2 GB · 7B ≈ 5 GB · 14B ≈ 9 GB · 32B ≈ 19 GB · 70B ≈ 40 GB. These figures increase slightly if you raise the context length beyond 4096.
- Reference GPUs
- RTX 3060 12GB comfortably covers 7B-14B. RTX 4070 12GB: same, with headroom. RTX 4080 16GB: 14B in Q5, the start of 24B. RTX 4090 24GB: 32B in Q4, 70B in Q3. Mac M4 Pro with 24-48GB unified memory: the same thing on the Apple Silicon side, but through the macOS version of LM Studio.
#Common pitfalls
- The AppImage refuses to start
- Run it from a terminal to see the error message. Most common: libfuse2 missing on Ubuntu 22.04+ (sudo apt install libfuse2 fixes the problem). Next: glibc is too old (LM Studio requires 2.35+, so no CentOS 7 or Debian 10).
- The GPU is not being used
- Check nvidia-smi (or rocm-smi). If the GPU isn’t listed: drivers are missing or too old. If it is listed but LM Studio remains on the CPU: Settings → Hardware lets you force the backend (CUDA/ROCm/Vulkan). On AMD, the user must be in the render and video groups—log out and back in after running usermod.
- The model is slow despite the GPU
- VRAM is saturated and partial offloading is happening silently. Reduce the context length, lower the GPU offload slider to measure the threshold, or use a more aggressive quantization (Q5 → Q4 → Q3).
- Port 1234 is already in use
- Often occupied by another development service. Change the port in Developer → Settings, or kill the process: ss -tulpn | grep 1234. There is no need to touch the firewall for localhost use.
- Breaking automatic updates
- LM Studio updates itself by default. In a controlled environment (enterprise, CI/CD integration), disable this via Settings → Updates. Note the version number that works so you can roll back if needed.
- No CLI autocompletion
- The AppImage does not provide a rich CLI. For serious scripting, the server’s HTTP API (port 1234) remains the cleanest option. If you need a real CLI, llama.cpp or Ollama are better suited.
#Go further
You have LM Studio installed, a GGUF loaded, and an OpenAI-compatible server running. The natural next directions are:
- Deep dive into the API server
- The Transformer LM Studio as an API server guide covers streaming, multiple models in parallel, secure LAN exposure, and performance tuning.
- Compare with Ollama on Linux
- The guide Installing Ollama on Linux covers the other popular stack, which is more daemon/server-oriented. For many users, the combination of Ollama (daemon) + LM Studio (client interface on top) is ideal.
- Choosing the right quantization
- The Choosing Your Quantization (Q4, Q5, Q8, FP16) guide visually compares actual quality loss and helps weigh quality against VRAM, especially when downloading your own GGUFs.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.