How to install Ollama: Step-by-step guide for beginners

You are trying to find out how to install Ollama to run LLM models locally on your machine? This guide is specifically designed for you, a Mac or PC user who wants to explore the world of open weights without relying on online services. Ollama dramatically simplifies this complex process. We'll detail every step you need to start interacting with powerful models today. This guide covers installation, the first run, and best practices for using models in your local environment.

What is Ollama, and why use it?

Ollama is an open-source tool that lets you download, configure, and run LLMs (Large Language Models) directly on your computer. Instead of having to manage complex dependencies such as PyTorch, CUDA, or specific quantization formats manually, Ollama provides a simple, standardized CLI.

For users who want to test the power of open-weight models without requiring a heavy server infrastructure, this is the ideal solution. You can experiment with advanced architectures such as MiMo V2.5 Pro (1020B) or lighter versions such as Mistral Small 4 (119B), while checking the power-consumption specifications required for your hardware see the official Ollama documentation.

The main advantage is simplicity: one command is often enough to download and run a model, making experimentation much more accessible to beginners. We'll see how this works in practice in the following sections. To deepen your knowledge of the available models, see our reference catalog.

Detailed installation guide for Mac and Windows

The process how to install Ollama varies slightly depending on your operating system, but the principle remains the same: download the appropriate executable and launch the service.

For macOS users (Apple Silicon)

  1. Go to Ollama's official page ici.
  2. Click the download link for macOS.
  3. Run the downloaded application. Ollama will install and launch in the background, automatically configuring the required paths.

For Windows users (via WSL or natively)

Although Ollama is often used through Linux environments, a Windows installation is available. Using the Windows Subsystem for Linux (WSL2) is strongly recommended to ensure the best compatibility with standard CLI commands and avoid GPU library path issues.

  1. Make sure WSL2 is enabled on your system.
  2. Download and install Ollama by following the Windows-specific instructions provided by the development team in their GitHub documentation.
  3. Once the service has started, open your WSL terminal to begin operations.

Technical note: Model efficiency depends heavily on hardware capacity. For example, running DeepSeek V4 Pro 1.6T (1600B) in Q4 requires a significant amount of VRAM, although its configuration can be optimized for different environments see the detailed specifications on our website. For more advanced technical information on local inference, consult resources such as those available on HuggingFace HuggingFace Documentation.

Getting Started: Running a Model from the Command Line

Once Ollama is installed and the service is active, you are ready to interact with open-source LLMs. The syntax is extremely simple. To test your installation, we will download and run a small model.

The basic command to how to install ollama ends with using a specific model. For example, to run the model llama3, you will use:

ollama run llama3

Ollama will check whether this model is available locally. If it is not, it will automatically download the default version (often an optimized quantization), then launch an interactive chat session with the LLM.

Choosing and testing high-performance models

quelllm.fr’s catalog offers hundreds of options to refine your choice based on your needs:

Before starting a large download, we recommend checking the model pages to verify the license (MIT, Apache 2.0, etc.) and estimated VRAM usage on our platform. If you are looking for information on developing these models, see their respective pages on GitHub GitHub LLM Repositories.

Optimizing and managing local LLM resources

One of the main challenges when using Ollama is managing hardware resources (VRAM and system RAM). Models are often quantized to reduce their memory footprint, but this affects output quality.

Understanding quantization formats

When you download a model through Ollama, it is generally provided in a quantized version (for example, Q4_K_M). Quantization reduces the number of bits used to represent each neural-network weight, drastically reducing the required VRAM. However, this can result in a slight loss of accuracy compared with the original FP16 or BF16 version.

For example, if you compare Nemotron 3 Ultra (550B) in Q4 (~319 GB of VRAM), you’ll get a good performance/size compromise. If you have a very powerful graphics card, testing the non-quantized or BF16 version may be useful for evaluating the model’s maximum potential, as with Nemotron 3 Ultra Base (BF16).

Concrete use cases and benchmarks

Ollama excels in the following use cases:

For a structured comparison of different models in terms of capacity and size, we invite you to consult our comparison tool. For an in-depth analysis of recent architectures, refer to publications on arXiv.

FAQ: Frequently asked questions about installing Ollama

Q: Do I need a very powerful graphics card to use Ollama?

R : It depends on the model you want to run. For smaller models (e.g.: Mixtral 8x22B Instruct, ~82 GB VRAM), a decent GPU may be enough. Models with several hundred billion parameters, such as DeepSeek V4 Pro 1.6T, will require a massive amount of VRAM or will have to run using system RAM, which is much slower.

Q: How can I tell which model to choose for my setup?

R : See our configurator on quelllm.fr. You will enter your machine specifications there (VRAM, RAM), and we will recommend optimized models, for example by comparing Llama 3.1 405B Instruct à Qwen 3.5 397B-A17B.

Q: Does Ollama work only with English-language models?

R : No, many open-weight LLMs are multilingual. Models like Kimi K2.6 or Ling 2.6 1T have demonstrated strong capabilities in French and other languages, although performance may vary depending on the specific model selected and its initial training.

Q: Is Ollama free?

R : Yes, Ollama itself is an open-source tool, and using it to run the models listed on our platform (which are under open licenses such as MIT or Apache 2.0) is free. Any costs are related to your machine's energy use during inference.

Q: How do I update my models after installation?

R : Updating is very simple from the command line. If you want a newer version of a model, use ollama pull nom_du_modele:version or simply ollama run nom_du_modele so that Ollama automatically checks for and downloads the latest available iterations from the community registry.

Conclusion: Your journey to a local LLM starts here

By following this guide, you know how to install Ollama on your Mac or PC to get started immediately with powerful local LLMs. Whether you want to test the robustness of GLM 5.2 753B-A40B or the efficiency of MiniMax M3, Ollama is the bridge between the model and your terminal. To take the next step and compare the technical performance (VRAM, tokens/sec) of the available models, visit our full catalog. Happy exploring!

The hardware for running an LLM locally

To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:

Amazon GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) →

Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.

Article published and updated on by Mohamed Meguedmi · Data source: /api/models.json · Content license: CC BY 4.0.

An error or update to report? Contribute.