Install an LLM locally: the step-by-step guide (2026)
Installing an LLM locally requires three things: a machine with enough RAM or VRAM for your target model size, a tool to run it — LM Studio without a terminal, Ollama from the command line, or llama.cpp for experts — and a first quantized model suited to that hardware. Allow 10 to 15 minutes to get a working first local chat on a recent machine, with no internet connection needed once the model has been downloaded.
Do you want to run AI directly on your PC or Mac, but all the tool and model names look alike and you don't know where to start? This guide provides the missing overview before you choose: verify that your machine can keep up, compare the three installation methods, identify your first model, and avoid the mistakes that ruin a first attempt. Each step links to a more detailed dedicated guide if you want to go further on a specific point.
#Check your machine: RAM, VRAM, and model size
Before choosing a tool, the question that really matters is: how much memory can your machine dedicate to a model? A “7B” or “32B” model refers to its number of parameters, but its quantized version — compressed to fit in memory, at the cost of a small loss in precision — determines the space actually required. Q4 quantization (often written Q4_K_M) is the standard compromise for getting started: it greatly reduces the memory required compared with the model’s native weights, with a quality loss that is generally barely noticeable in use.
- 8 GB of RAM, no dedicated GPU
- lightweight models around 3-4B in Q4: basic use, slower but functional responses.
- 16 GB of RAM (CPU) or 8 GB of VRAM (GPU)
- the most comfortable entry point, with 7-8B models in Q4—the most common format for a first serious use case.
- 12 GB of VRAM
- enough headroom to move up to a 14B model without any particular difficulty.
- 24 GB of VRAM (or 32–48 GB of unified memory on Mac)
- a 32B model with Q4 quantization runs comfortably.
- 64 GB or more of RAM or unified memory
- models around 70B become viable, with partial offloading between the CPU and GPU and significantly lower throughput.
#Choose your method: LM Studio, Ollama, or llama.cpp
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
There are three main ways to run an LLM locally. They are not mutually exclusive, but they start from different philosophies—the choice depends mainly on your comfort level with a terminal and what you plan to do with the model once it is installed.
- LM Studio — zero terminal
- Desktop application with a full graphical interface, integrated model search, and visual GPU settings. The most straightforward choice for a first experience, available on Windows, macOS (Apple Silicon), and Linux.
- Ollama — terminal and API
- one-command installation, a command-line-driven model library, and a local OpenAI-compatible API server active by default. The natural choice if you plan to script, automate, or connect a third-party tool.
- llama.cpp — for experts
- the inference engine that many other tools use behind the scenes, including LM Studio. Compile from source, control every parameter precisely, no interface — for those who want to understand or optimize every detail.
In short: you don't want to open a terminal → LM Studio. You want to call your model from a script, automation, or application → Ollama. You want to understand or adjust every compilation parameter, or run the model on unusual hardware → llama.cpp.
#Express installation, step by step
Here is the condensed installation for each of the three paths. The goal is to provide an overview so you can get started in a few minutes; a dedicated guide exists for each tool if you want the full details, including screenshots.
With LM Studio (zero terminal):
- 011. Download the installerFrom the official LM Studio website, download the version corresponding to your system.
- 022. Open the model searchOn first launch, open the search tab (Ctrl/Cmd+Shift+M shortcut) to browse the catalog.
- 033. Choose a suitable modelThe app displays a VRAM compatibility estimate before downloading: prioritize a model marked as compatible with your machine.
- 044. Load the model and chatOpen the Chat tab, load the downloaded model from the top menu, and ask your first question.
With Ollama (terminal and API):
- 011. Install OllamaOfficial script for Linux, .exe or .dmg installer for Windows and macOS. The desktop application launches automatically after installation.
- 022. Run your first modelIn a terminal, the ollama run suivie command with a model name downloads it and then starts a chat directly in the terminal.
- 033. Discuss or connect a third-party toolContinue in the terminal, or point an OpenAI-compatible application to the local API server (port 11434 by default).
- 044. Manage your installed modelsList, change, or delete your models at any time using the dedicated management commands.
With llama.cpp (experts):
- 011. Compile the binaryClone the repository and compile it for your hardware (CPU, CUDA, Metal, or ROCm, depending on your configuration).
- 022. Download a GGUF modelRetrieve a file in GGUF format from Hugging Face, choosing the quantization suited to your available memory.
- 033. Launch the chat binaryPoint it to the downloaded GGUF file, with your preferred context and GPU offloading parameters.
- 044. Adjust GPU/CPU offloadingRefine the GPU/CPU split layer by layer to find the best speed/memory trade-off for your setup.
#Download your first model
The right first model depends mainly on what your machine can comfortably keep in memory, not on the model's reputation. Here are reliable guidelines for your configuration.
- Modest machine (8-16 GB of RAM, no dedicated GPU)
- a small general-purpose model around 3–4B in Q4: fast even on pure CPU, and more than sufficient for chatting, summarizing, or translating a short text.
- 8-12 GB VRAM
- a 7–8B model in Q4_K_M, such as Qwen3 8B or Mistral: the most common quality/speed tradeoff for daily use.
- 16–24 GB of VRAM
- a 14B model, or a 32B in Q4 if you're at the top of this range, such as Gemma in its largest version: more headroom for reasoning or code.
- You don't yet know what you want to use it for
- start small. Verifying that a lightweight model responds correctly takes a few minutes; you can move up in size once the first test succeeds.
#Common mistakes to avoid
Most bad first impressions of local AI come from poor configuration rather than a real problem with the tool or model. Here are the most common pitfalls for people installing their first local LLM, and how to spot them quickly.
- Model too large for VRAM
- The model spills into system RAM or even crashes during loading. Responses become extremely slow, or an out-of-memory error appears. Step down one model size or one quantization level.
- Randomly chosen quantization
- Downloading the heaviest available version “to get the best” often ends up exceeding available memory. Start with Q4_K_M, and move up only if the memory headroom truly allows it.
- Fan revving up and sluggish responses
- Normal for a first large model: inference heavily taxes the GPU or CPU. If it is too slow for your taste, that indicates the model is oversized for your machine, not that there is a bug to fix.
- Download everything before testing
- It is better to verify that a small model works and responds correctly before starting the download of a model several tens of GB in size.
#What next? RAG, coding copilot, API
Once you have a first model installed and working, several natural next steps open up depending on your use case: chat with your own documents, connect a code copilot to your editor, or expose the local API server to your own scripts and applications.
- Chat with your documents (RAG)
- LM Studio offers built-in RAG out of the box. With Ollama, you need to pair it with a third-party tool such as Open WebUI or a dedicated pipeline to achieve an equivalent result.
- Copilot your code editor
- Ollama's or LM Studio's OpenAI-compatible API server can be connected to extensions such as Cline for coding with 100% local assistance.
- Automate via the API
- Both tools expose an OpenAI-compatible local server that can be reused in any script or application already designed for this API, with no client-side changes.
#Frequently asked questions
Is installing an LLM locally legal and free?+
What is the minimum configuration to get started?+
Can you install an LLM locally without a graphics card?+
How much disk space should you plan for?+
Which model should you choose to get started?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.