Beginner 12 minInstallation

Install an LLM locally: the step-by-step guide (2026)

Direct response

Installing an LLM locally requires three things: a machine with enough RAM or VRAM for your target model size, a tool to run it — LM Studio without a terminal, Ollama from the command line, or llama.cpp for experts — and a first quantized model suited to that hardware. Allow 10 to 15 minutes to get a working first local chat on a recent machine, with no internet connection needed once the model has been downloaded.

Do you want to run AI directly on your PC or Mac, but all the tool and model names look alike and you don't know where to start? This guide provides the missing overview before you choose: verify that your machine can keep up, compare the three installation methods, identify your first model, and avoid the mistakes that ruin a first attempt. Each step links to a more detailed dedicated guide if you want to go further on a specific point.

By Mohamed Meguedmi·Update 2026-08-24·Tested on Windows, macOS, and Linux

#Check your machine: RAM, VRAM, and model size

Before choosing a tool, the question that really matters is: how much memory can your machine dedicate to a model? A “7B” or “32B” model refers to its number of parameters, but its quantized version — compressed to fit in memory, at the cost of a small loss in precision — determines the space actually required. Q4 quantization (often written Q4_K_M) is the standard compromise for getting started: it greatly reduces the memory required compared with the model’s native weights, with a quality loss that is generally barely noticeable in use.

8 GB of RAM, no dedicated GPU
lightweight models around 3-4B in Q4: basic use, slower but functional responses.
16 GB of RAM (CPU) or 8 GB of VRAM (GPU)
the most comfortable entry point, with 7-8B models in Q4—the most common format for a first serious use case.
12 GB of VRAM
enough headroom to move up to a 14B model without any particular difficulty.
24 GB of VRAM (or 32–48 GB of unified memory on Mac)
a 32B model with Q4 quantization runs comfortably.
64 GB or more of RAM or unified memory
models around 70B become viable, with partial offloading between the CPU and GPU and significantly lower throughput.
i
These are reference points, not guarantees
The context length used, the operating system, and other open applications change the actual available headroom. Always keep a margin of about 20 to 30% beyond the model’s strict size before considering a configuration comfortable.

#Choose your method: LM Studio, Ollama, or llama.cpp

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

There are three main ways to run an LLM locally. They are not mutually exclusive, but they start from different philosophies—the choice depends mainly on your comfort level with a terminal and what you plan to do with the model once it is installed.

LM Studio — zero terminal
Desktop application with a full graphical interface, integrated model search, and visual GPU settings. The most straightforward choice for a first experience, available on Windows, macOS (Apple Silicon), and Linux.
Ollama — terminal and API
one-command installation, a command-line-driven model library, and a local OpenAI-compatible API server active by default. The natural choice if you plan to script, automate, or connect a third-party tool.
llama.cpp — for experts
the inference engine that many other tools use behind the scenes, including LM Studio. Compile from source, control every parameter precisely, no interface — for those who want to understand or optimize every detail.

In short: you don't want to open a terminal → LM Studio. You want to call your model from a script, automation, or application → Ollama. You want to understand or adjust every compilation parameter, or run the model on unusual hardware → llama.cpp.


#Express installation, step by step

Here is the condensed installation for each of the three paths. The goal is to provide an overview so you can get started in a few minutes; a dedicated guide exists for each tool if you want the full details, including screenshots.

With LM Studio (zero terminal):

  1. 01
    1. Download the installer
    From the official LM Studio website, download the version corresponding to your system.
  2. 02
    2. Open the model search
    On first launch, open the search tab (Ctrl/Cmd+Shift+M shortcut) to browse the catalog.
  3. 03
    3. Choose a suitable model
    The app displays a VRAM compatibility estimate before downloading: prioritize a model marked as compatible with your machine.
  4. 04
    4. Load the model and chat
    Open the Chat tab, load the downloaded model from the top menu, and ask your first question.

With Ollama (terminal and API):

  1. 01
    1. Install Ollama
    Official script for Linux, .exe or .dmg installer for Windows and macOS. The desktop application launches automatically after installation.
  2. 02
    2. Run your first model
    In a terminal, the ollama run suivie command with a model name downloads it and then starts a chat directly in the terminal.
  3. 03
    3. Discuss or connect a third-party tool
    Continue in the terminal, or point an OpenAI-compatible application to the local API server (port 11434 by default).
  4. 04
    4. Manage your installed models
    List, change, or delete your models at any time using the dedicated management commands.

With llama.cpp (experts):

  1. 01
    1. Compile the binary
    Clone the repository and compile it for your hardware (CPU, CUDA, Metal, or ROCm, depending on your configuration).
  2. 02
    2. Download a GGUF model
    Retrieve a file in GGUF format from Hugging Face, choosing the quantization suited to your available memory.
  3. 03
    3. Launch the chat binary
    Point it to the downloaded GGUF file, with your preferred context and GPU offloading parameters.
  4. 04
    4. Adjust GPU/CPU offloading
    Refine the GPU/CPU split layer by layer to find the best speed/memory trade-off for your setup.
i
For the full details
Each of these three paths deserves its own step-by-step guide, with screenshots and advanced options. This condensed version is enough to get a first model responding; return to the dedicated guide when you want to explore a specific topic, such as AMD GPU support or the network configuration of a server.

#Download your first model

The right first model depends mainly on what your machine can comfortably keep in memory, not on the model's reputation. Here are reliable guidelines for your configuration.

Modest machine (8-16 GB of RAM, no dedicated GPU)
a small general-purpose model around 3–4B in Q4: fast even on pure CPU, and more than sufficient for chatting, summarizing, or translating a short text.
8-12 GB VRAM
a 7–8B model in Q4_K_M, such as Qwen3 8B or Mistral: the most common quality/speed tradeoff for daily use.
16–24 GB of VRAM
a 14B model, or a 32B in Q4 if you're at the top of this range, such as Gemma in its largest version: more headroom for reasoning or code.
You don't yet know what you want to use it for
start small. Verifying that a lightweight model responds correctly takes a few minutes; you can move up in size once the first test succeeds.
→
The quantization name is not cosmetic
A tag such as Q4_K_M, Q5_K_M, or Q8_0 indicates the model’s compression level. The lower the number, the lighter and faster the model, but the less accurate it is. Q4_K_M remains the best starting point for a first try.

#Common mistakes to avoid

Most bad first impressions of local AI come from poor configuration rather than a real problem with the tool or model. Here are the most common pitfalls for people installing their first local LLM, and how to spot them quickly.

Model too large for VRAM
The model spills into system RAM or even crashes during loading. Responses become extremely slow, or an out-of-memory error appears. Step down one model size or one quantization level.
Randomly chosen quantization
Downloading the heaviest available version “to get the best” often ends up exceeding available memory. Start with Q4_K_M, and move up only if the memory headroom truly allows it.
Fan revving up and sluggish responses
Normal for a first large model: inference heavily taxes the GPU or CPU. If it is too slow for your taste, that indicates the model is oversized for your machine, not that there is a bug to fix.
Download everything before testing
It is better to verify that a small model works and responds correctly before starting the download of a model several tens of GB in size.

#What next? RAG, coding copilot, API

Once you have a first model installed and working, several natural next steps open up depending on your use case: chat with your own documents, connect a code copilot to your editor, or expose the local API server to your own scripts and applications.

Chat with your documents (RAG)
LM Studio offers built-in RAG out of the box. With Ollama, you need to pair it with a third-party tool such as Open WebUI or a dedicated pipeline to achieve an equivalent result.
Copilot your code editor
Ollama's or LM Studio's OpenAI-compatible API server can be connected to extensions such as Cline for coding with 100% local assistance.
Automate via the API
Both tools expose an OpenAI-compatible local server that can be reused in any script or application already designed for this API, with no client-side changes.
→
One step at a time
There is no need to target RAG, a coding copilot, and automation on day one. First validate a local chat that answers your questions correctly; more advanced use cases naturally follow once this foundation is solid.

#Frequently asked questions

Is installing an LLM locally legal and free?+
Yes: tools such as Ollama, LM Studio, or llama.cpp are free, and open-weight models (Llama, Qwen, Mistral, Gemma...) are released under licenses that permit personal use. Some licenses have specific conditions for large-scale commercial use, which you should verify case by case on the model page.
What is the minimum configuration to get started?+
In practice, 16 GB of RAM and a recent processor are enough for a first model with a few billion parameters. A GPU with at least 8 GB of VRAM, or a Apple Silicon chip with enough unified memory, makes the experience much more comfortable as you move to larger models.
Can you install an LLM locally without a graphics card?+
Yes, all three methods presented here run entirely on the CPU. Models from 3B to 8B remain usable without a GPU, but responses are slower than with a dedicated card.
How much disk space should you plan for?+
Each quantized model ranges from a few hundred MB for the smallest models to several dozen GB for the largest. Leave plenty of free space on an SSD if you plan to test several models before settling on one.
Which model should you choose to get started?+
A general-purpose 7 to 8B model, such as Qwen3 8B or Mistral, in Q4_K_M quantization, is a solid starting point on most recent machines. Then adjust the size based on your actual RAM or VRAM and your initial tests.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.