Complete LM Studio tutorial: Install and run an LLM

If you are looking for a lm studio tutorial to get started with open-source models locally, this guide is for you. LM Studio has established itself as the go-to tool for exploring the potential of large language models (LLMs) directly on your machine, whether you're using Windows, macOS, or Linux. We'll walk through every step, from installation to your first request with high-performance models available in our catalog catalog. This complete tutorial will cover the technical aspects for optimal use of your local LLM.

🚀 Initial Installation and Configuration of LM Studio

LM Studio is designed to simplify the interface between user hardware and the complexity of model weights. Before diving into launching it, it's crucial to install the application correctly. For users on macOS- or Windows-based systems, the process is generally straightforward through their official website. If you're on Linux, see our dedicated guide lm studio linux installation guide.

Once the application is launched, the first step is to navigate to the search section (often represented by a magnifying glass) to find the desired model. We recommend consulting our comparator to identify the best candidates based on your hardware configuration and the LLM's requirements.

Choosing the right model is critical. For example, if you have substantial VRAM, you might consider massive models such as Kimi K3 (2800B) or DeepSeek V4 Pro 1.6T (1600B), although deploying them requires robust infrastructure. To get started gradually, medium-sized models are ideal for testing features without overloading your GPU.

🧠 Choose and Download an Optimized LLM

The core of the process is downloading the model weights. LM Studio often uses quantized formats (such as Q4_K_M) that drastically reduce the required memory footprint, allowing very large models to run on consumer graphics cards.

When you select a model on our platform Hugging Face Hub, make sure to check the technical specifications: * Size (Parameters) : Determines the model's theoretical capacity. * Quantization : The compression level (Q4, Q5, etc.) directly affects VRAM usage and response quality. * License : Check whether the LLM is compatible with your commercial or personal use (e.g., Apache 2.0 for Mistral Large 3 675B).

For example, if you're targeting good coding performance, models like Kimi K2.7 Code (1059B) are relevant. To assess feasibility on your machine, we recommend consulting our configurator.

⚙️ The Local Loading and Inference Process

Once the LLM file has been downloaded into LM Studio (it is generally stored locally), you move on to inference. You navigate to the chat interface or the API Server section, depending on your needs.

Loading the model into GPU memory depends heavily on your graphics card and the amount of available VRAM. For models such as DeepSeek V4 Pro 0813 1.7T, a configuration with approximately 986 GB of VRAM (in Q4) is required, which is an extreme case. Conversely, for quick testing, you can load lighter models such as Mixtral 8x22B Instruct (approximately 82 GB in Q4).

If you want to integrate your local LLM into other applications or create your own agents, it is recommended that you configure LM Studio as an API server. This mode allows tools such as Open WebUI to communicate with your locally hosted model open-webui ollama tutorial.

🛠️ Advanced Use Cases: RAG and Local API

LM Studio is more than just a conversation. It can serve as an engine for more complex applications.

1. Retrieval-Augmented Generation (RAG) : For your LLM to respond based on your private documents, you’ll use built-in RAG features or plugins. This allows a model like Inkling (975B) to access a specific knowledge base without having been trained on it. See our lm studio rag local documents.

2. Local API Server : By enabling server mode, you expose the LLM's capabilities to standard HTTP requests. This is essential for integration into automated workflows or with frameworks such as LangChain local Python LangChain AI agent.

To compare this local approach with other solutions, we also offer comparative guides between LM Studio and Ollama: ollama vs lm studio which one to choose.

❓ FAQ on Using LM Studio

Q: What is the best model for getting started with LM Studio?

For a first experience without requiring extreme resources, we suggest trying Qwen 3.5 122B-A10B (73 GB in Q4). It offers a good balance between performance and hardware requirements, while also benefiting from the Apache 2.0 license Alibaba on Hugging Face.

Q: How do I know whether my PC can run an LLM?

VRAM is the critical factor. Use our configurator to estimate requirements. For example, GLM 5.3 Flash 320B-A18B requires approximately 186 GB in Q4, which is demanding.

Q: What is the difference between using LM Studio and Ollama?

LM Studio provides a full graphical interface for model management and local experimentation. Ollama is a lighter CLI tool, ideal for quickly integrating into scripts how to install ollama.

Q: Are performance figures (tokens/sec) stable?

Tokens/sec depends on your hardware and the model you choose. For a realistic estimate, you can check the available benchmarks on the Open LLM Leaderboard.

🏁 Conclusion: Getting Started Locally with LM Studio

Ce lm studio tutorial guided you through installation, weight selection, and local inference methods to harness the power of open-source LLMs. Whether you want to experiment with a cutting-edge model like DeepSeek V4 Flash 284B or simply test a small architecture, LM Studio is your gateway to local execution. To refine your choices based on your hardware, visit our configurator or explore our catalog.

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.