How to install Ollama: Complete guide (Windows, Mac, Linux)
Are you looking to run open-source LLMs locally? Learn how to install ollama is the first step toward democratizing access to open weights on your own machine. This tool greatly simplifies downloading and running various models, whether you're on Windows, macOS, or Linux. This complete guide details each installation method so you can start experimenting with powerful architectures such as DeepSeek V4 Pro 1.6T or MiMo V2.5 Pro. We will explore the technical steps, performance, and how to leverage this local platform.
What is Ollama, and why use it for local LLMs?
Ollama is a tool designed to simplify the deployment of large language models (LLMs) on personal environments or local servers. Instead of manually managing complex dependencies, quantization frameworks such as llama.cpp, and the configurations specific to each model, Ollama provides a simple CLI interface. It acts as a local server that lets you interact with different LLMs through a standardized API.
For those who want to take local experimentation further, it is useful to know that tools such as llama.cpp (official GitHub) are often the foundation under the hood of Ollama. Whether you want to test a high-performance model like DeepSeek V4 Flash 0731 304B or explore the capabilities of Kimi K2.5, Ollama makes access to these weights more direct.
The main advantage is simplicity: one command is enough to download and run a model, which is perfect for users who want to move quickly from concept to local execution on their PC or Mac. If you want to compare different architectures before choosing your setup, our catalog lists hundreds of models with their respective VRAM requirements and licenses. You can review the work of DeepSeek on Hugging Face to see the range of their creations available locally.
Detailed installation guide for each operating system
The process varies slightly by OS, but the principle remains the same: install the Ollama executable and then use the command line to run the models.
Installation on macOS (Mac)
For Mac users, whether using M-series chips or Intel machines, installation is particularly smooth. You can download the official installation script from Ollama (official GitHub). Once installed, Ollama runs in the background and you are ready to interact with the models through your terminal. For specific Mac configurations, particularly to optimize use of the Apple Silicon GPU, see our guide llm-mac-mini-m4.
Installation on Windows
On Windows, you have two main options: use the native binary or go through WSL2 (Windows Subsystem for Linux 2). If you choose native installation, follow the instructions on the official website. For more robust integration and better compatibility with Linux-based development tools, we often recommend using WSL2. Our guide ollama-wsl2-vs-windows-natif explains this approach in detail.
Installation on Linux
Installation on Linux is generally most straightforward via a repository script or a precompiled binary, following the instructions provided by Ollama (official GitHub). For step-by-step documentation specific to shell commands, refer to our dedicated tutorial: installer-ollama-linux.
Run and test your first LLM models with Ollama
Once Ollama is installed, you use the command ollama run <nom_du_modele> to get started. This is where you choose the brain for your local session.
For example, if you want to experiment with a highly capable model like DeepSeek V4 Pro 0813 1.7T, you will run the corresponding command. It's important to note that models are often offered in different versions (quantizations) to balance size and performance. For example, models such as GLM 5.2 753B-A40B will require substantial hardware.
To assess raw power, you can consult the Open LLM Leaderboard (Hugging Face). If your goal is to set up a complete environment for agent development, we’ve prepared tutorials on integrating with CrewAI or local agents crewai-ollama-multi-agents.
Model examples to try: * For advanced reasoning capabilities, try Inkling (975B) or DeepSeek V4 Pro 1.6T (1600B). You can also look at the capabilities of Kimi K3. * If you're interested in code, Kimi K2.7 Code is a good starting point for testing programming-specialized models. * For lighter and faster models locally, look at variants such as Mixtral 8x22B Instruct.
Optimizing and making advanced use of Ollama
Using local LLMs is not limited to simple terminal conversations. Ollama can be integrated into more complex workflows. If you want a user-friendly graphical interface, we recommend exploring tools such as Open WebUI (official GitHub), which integrates seamlessly with server Ollama.
For advanced users who want to automate tasks or build applications, you can configure direct API calls to your Ollama instance. Our guide appel-outil-ollama-tutoriel shows you how to integrate these LLMs into Python scripts.
If you encounter performance issues, particularly related to GPU usage (GPU acceleration), see our guide depanner-ollama-gpu-erreurs. For a direct comparison between Ollama and other solutions such as LM Studio, see our comparison in the guides interfaces-graphiques-ollama-comparatif. If you're looking for more specific alternatives, our page alternatives-ollama-tour-horizon is a useful resource.
FAQ: Frequently asked questions about installing Ollama
Q: What is the minimum hardware requirement to get started with Ollama?
A: For basic tests, a modern system is sufficient. However, if you want to run larger models such as GLM 5.3 Flash 320B-A18B (which requires approximately 186 GB in Q4), a graphics card with substantial VRAM is necessary. For modest configurations, choose quantized models or use the CPU only if you are patient during inference.
Q: How do you choose a model after installing Ollama?
A: The choice depends on your use case (chat, code, reasoning) and hardware resources. If speed is key, look at “Flash” versions such as DeepSeek V4 Flash 284B. For maximum quality on complex tasks, explore larger models such as Kimi K3 or DeepSeek V4 Pro 0813 1.7T, always verifying their specifications in our catalog.
Q: Does Ollama work only with GGUF models?
A: Although Ollama widely uses the GGUF format (optimized for CPU/GPU inference via llama.cpp), it also supports other formats and can interact with weights from various sources, following conventions defined by the open-source LLM community Hugging Face Hub (documentation).
Q: Can I use Ollama for RAG?
A: Yes, absolutely. Although Ollama is the inference engine, you need to integrate it into a RAG (Retrieval Augmented Generation) framework such as LangChain or LlamaIndex. We have specific tutorials on this topic, for example anythingllm-rag-tutoriel-complet.
Q: Is Ollama free and open source?
A: The tool itself, Ollama, is designed to be accessible and easy to use. The models it runs mostly use open licenses (MIT, Apache 2.0) or follow the vendors' open policies, such as Moonshot AI on Hugging Face, which guarantees free local use for experimentation.
Conclusion: Your gateway to local LLMs
In summary, how to install Ollama is a simple but powerful process that unlocks access to cutting-edge models on your own infrastructure. Whether you are a beginner or an expert looking to integrate complex systems such as those based on Llama 3.1 405B Instruct, Ollama is the ideal orchestration solution. Once installation is complete, explore our configurator to find the perfect model for your hardware and specific needs.
The hardware for running an LLM locally
To run these models comfortably locally, a RTX 5070 Ti offers an excellent price/performance ratio. Compare prices:
Affiliate links — QuelLLM may earn a commission on purchases, at no extra cost to you. As an Amazon Associate, BestLLMfor earns from qualifying purchases.