CUDA: what is it, and do you need to install it for AI ?
CUDA is the proprietary software platform that NVIDIA launched in 2007 to use its GPUs for general-purpose computing, not just graphics. It works only on NVIDIA hardware and underpins nearly all AI software. For inference with Ollama or LM Studio, an up-to-date NVIDIA driver is sufficient: the CUDA libraries are already bundled. The complete Toolkit is needed only to compile code from source.
CUDA is NVIDIA’s software platform that lets you use a graphics card for general-purpose computing, not just displaying images. Launched in 2007, it only works on NVIDIA hardware. Almost all AI software targets it first, which explains why GeForce cards are the easiest path to local AI. As of September 20, 2026, one misconception remains widespread: to run Ollama or LM Studio, you do not need to “install CUDA.” This page explains what CUDA is, what you actually need to install, and how to read version numbers.
#What is CUDA?
A graphics card contains thousands of small processors designed to apply the same operation to very large amounts of data at the same time, rather than performing one complex operation at a time like a CPU. Originally, they could only be used through graphics interfaces: you had to disguise a computation as an image to send it through the 3D pipeline. CUDA, short for Compute Unified Device Architecture, eliminated that disguise when it launched in 2007. A programmer writes ordinary code in C or C++, or calls it from Python, and it runs directly on those processors without going through graphics rendering.
In everyday usage, “CUDA” refers to four things at once: the programming model, the nvcc compiler, the part of the driver that executes the code, and a set of optimized libraries such as cuBLAS for linear algebra and cuDNN for neural networks. Frameworks such as PyTorch rely on these libraries, and AI applications in turn rely on the frameworks: it is a stack of successive layers, with each level hiding the complexity of the one below it. When a tool advertises “CUDA compatibility,” this concretely means that its heavy computations—especially matrix multiplications and convolutions—ultimately run in these proprietary NVIDIA libraries rather than in generic code.
#What about CUDA cores?
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
This is the marketing name NVIDIA gives to a GPU’s parallel processors, its basic compute units. One RTX 4060 has 3,072, while one RTX 5090 has 21,760—almost seven times as many, according to the manufacturer’s specifications. More cores mean more calculations per second, which directly matters for image generation, model training, and an LLM’s prompt processing—the phase when the GPU must process all the text written so far at once.
This matters much less for what we notice most when using an LLM day to day: writing speed, token by token. There, the limit is how quickly the weights can be read from memory, not how quickly they can be multiplied once read. In practice, a card with fewer cores but more and faster VRAM will therefore often write faster than a card that is more powerful on paper but has narrower memory bandwidth.
#Driver, libraries, Toolkit: what needs to be installed?
| Component | What it is | Who needs it |
|---|---|---|
| NVIDIA driver | Bridges the system and the card; contains the CUDA runtime | Anyone with a NVIDIA card |
| Runtime libraries | cuBLAS, cuDNN, and the rest | Already included in most applications and PyTorch packages: you almost never install them yourself |
| CUDA Toolkit | The nvcc compiler, headers, and profiling tools | Only those who compile CUDA code, such as llama.cpp from source |
#Read its version numbers
These two numbers legitimately differ, and the first one fools almost everyone. When nvidia-smi displays « CUDA Version: 13.0 », it means your driver can run software compiled for CUDA up to version 13.0. It does not mean the Toolkit is installed. And software compiled for an older CUDA version works without problems with a recent driver.
#Computing capacity, generation by generation
Each GPU NVIDIA has a “compute capability” number that identifies its architecture and the instructions it natively understands. Software is compiled for specific capabilities, generally a range of versions covered by the same package, and projects regularly drop support for older ones as updates roll out. This, much more than raw compute speed, is what ends a card's career in local AI: a RTX 2060 remains physically functional for years after purchase, but some newer tools refuse to install on it because no code was compiled for its generation.
| Generation | Cards | Compute capacity | Local AI landscape in 2026 |
|---|---|---|---|
| Pascal | GTX 1060, 1070, 1080 Ti | 6.1 | Runs tools based on llama.cpp; no FP16 acceleration; abandoned by some recent frameworks |
| Turing | GTX 1660, RTX 2060 at 2080 Ti | 7.5 | Minimum required by vLLM (7.0); well supported |
| Ampere | RTX 3050 at 3090 Ti | 8.6 | Fully supported; adds BF16 |
| Ada Lovelace | RTX 4060 to 4090 | 8.9 | Fully supported; adds FP8 support |
| Blackwell | RTX 5050 to 5090 | 12.0 | Fully supported by recent software; older versions must be updated or recompiled; brings FP4 |
The last line explains the wave of “my RTX 50 isn't detected” messages: a binary compiled before this architecture existed contains no code for compute capability 12.0. The solution is a newer version of the tool, or PyTorch built for a recent version of CUDA.
#CUDA versus ROCm, Metal, and Vulkan
| Platform | Manufacturer | Hardware | Status of local LLMs |
|---|---|---|---|
| CUDA | NVIDIA | GPU NVIDIA | The reference target: it supports everything first |
| ROCm | AMD | Recent Radeon and Instinct | Good on Linux with supported cards; narrower card list |
| Metal | Apple | Apple Silicon | Excellent with llama.cpp and MLX |
| Vulkan | Khronos, open standard | Almost all GPUs, including integrated GPUs | Works almost everywhere with llama.cpp; generally slower than the native solution |
| SYCL / oneAPI | Intel | Intel Arc and integrated GPUs | Functional, smallest ecosystem |
CUDA's advantage isn't that other platforms are incapable of doing the calculations. It comes from nearly twenty years of libraries, tools, and proven code, built GPU by GPU since 2007 according to the project's official page. For inference with common tools, the gap has narrowed considerably: llama.cpp runs on all five platforms, with similar performance at the same quantization level on comparable hardware. For training, fine-tuning, and research code just published on GitHub, however, CUDA remains the only path that works from day one: it is almost always the first target, and sometimes the only one, for repositories accompanying a new scientific paper, often several months before a port to other platforms appears—if it ever does.
- Ollama on AMD GPUs with ROCm
- llama.cpp with Vulkan, the universal backend
- On Mac: MLX or llama.cpp?
- Source: Wikipedia, CUDA history and architecture
- Source: CUDA compute capability table by GPU (GPUSmith)
#What this changes when buying
- You want everything to work without troubleshooting
- Choose a NVIDIA card, and prioritize VRAM capacity over the number of CUDA cores shown on the box: VRAM determines whether a model loads, not raw compute power.
- You mainly run GGUFs in Ollama or LM Studio
- A Mac with Apple Silicon, or a recent Radeon under Linux with ROCm, is a legitimate choice: llama.cpp supports both platforms properly, and you lose little in day-to-day usability.
- You plan to fine-tune, use vLLM, or follow research repositories
- Not having CUDA will cost you time every week, between dependencies that fail to compile and notebooks published assuming a GPU NVIDIA is available by default.
- You are buying a used card
- Avoid anything below compute capability 7.5, regardless of the listed price: beyond being slow, some recent frameworks simply refuse to install.
#FAQ
What does CUDA mean?+
Do you need to install CUDA to run AI locally?+
Does CUDA work on an AMD or Intel card?+
Why don’t nvidia-smi and nvcc show the same version?+
Are more CUDA cores better for an LLM?+
RTX 50 and Blackwell data center cards: do they have the same computing capability?+
Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — recent NVIDIA card with 16 GB of VRAM, supported by CUDA. All AI hardware →
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.