Beginner 10 minConcepts

CUDA: what is it, and do you need to install it for AI ?

Direct response

CUDA is the proprietary software platform that NVIDIA launched in 2007 to use its GPUs for general-purpose computing, not just graphics. It works only on NVIDIA hardware and underpins nearly all AI software. For inference with Ollama or LM Studio, an up-to-date NVIDIA driver is sufficient: the CUDA libraries are already bundled. The complete Toolkit is needed only to compile code from source.

CUDA is NVIDIA’s software platform that lets you use a graphics card for general-purpose computing, not just displaying images. Launched in 2007, it only works on NVIDIA hardware. Almost all AI software targets it first, which explains why GeForce cards are the easiest path to local AI. As of September 20, 2026, one misconception remains widespread: to run Ollama or LM Studio, you do not need to “install CUDA.” This page explains what CUDA is, what you actually need to install, and how to read version numbers.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What is CUDA?

A graphics card contains thousands of small processors designed to apply the same operation to very large amounts of data at the same time, rather than performing one complex operation at a time like a CPU. Originally, they could only be used through graphics interfaces: you had to disguise a computation as an image to send it through the 3D pipeline. CUDA, short for Compute Unified Device Architecture, eliminated that disguise when it launched in 2007. A programmer writes ordinary code in C or C++, or calls it from Python, and it runs directly on those processors without going through graphics rendering.

In everyday usage, “CUDA” refers to four things at once: the programming model, the nvcc compiler, the part of the driver that executes the code, and a set of optimized libraries such as cuBLAS for linear algebra and cuDNN for neural networks. Frameworks such as PyTorch rely on these libraries, and AI applications in turn rely on the frameworks: it is a stack of successive layers, with each level hiding the complexity of the one below it. When a tool advertises “CUDA compatibility,” this concretely means that its heavy computations—especially matrix multiplications and convolutions—ultimately run in these proprietary NVIDIA libraries rather than in generic code.

#What about CUDA cores?

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

This is the marketing name NVIDIA gives to a GPU’s parallel processors, its basic compute units. One RTX 4060 has 3,072, while one RTX 5090 has 21,760—almost seven times as many, according to the manufacturer’s specifications. More cores mean more calculations per second, which directly matters for image generation, model training, and an LLM’s prompt processing—the phase when the GPU must process all the text written so far at once.

This matters much less for what we notice most when using an LLM day to day: writing speed, token by token. There, the limit is how quickly the weights can be read from memory, not how quickly they can be multiplied once read. In practice, a card with fewer cores but more and faster VRAM will therefore often write faster than a card that is more powerful on paper but has narrower memory bandwidth.

#Driver, libraries, Toolkit: what needs to be installed?

ComponentWhat it isWho needs it
NVIDIA driverBridges the system and the card; contains the CUDA runtimeAnyone with a NVIDIA card
Runtime librariescuBLAS, cuDNN, and the restAlready included in most applications and PyTorch packages: you almost never install them yourself
CUDA ToolkitThe nvcc compiler, headers, and profiling toolsOnly those who compile CUDA code, such as llama.cpp from source
→
In plain English
For Ollama, LM Studio, ComfyUI Desktop, or a simple pip install torch: update the driver, and that's it. Installing the multi-gigabyte Toolkit “just in case” is the most common unnecessary step in local AI tutorials.

#Read its version numbers

Two commands, two different kinds of information
nvidia-smi        # version du pilote, et version MAXIMALE de CUDA qu'il sait exécuter
nvcc --version    # version du Toolkit, s'il est installé

These two numbers legitimately differ, and the first one fools almost everyone. When nvidia-smi displays « CUDA Version: 13.0 », it means your driver can run software compiled for CUDA up to version 13.0. It does not mean the Toolkit is installed. And software compiled for an older CUDA version works without problems with a recent driver.

i
CUDA 13 changed one habit
For CUDA 13.x, NVIDIA requires driver version 580 or later, according to its official release notes. And since version 13.0, the Windows display driver has no longer been included with the Toolkit package: you now have to install it separately from the NVIDIA website. This is a common source of compilation failures for users who reuse an old Toolkit package downloaded before the change.

#Computing capacity, generation by generation

Each GPU NVIDIA has a “compute capability” number that identifies its architecture and the instructions it natively understands. Software is compiled for specific capabilities, generally a range of versions covered by the same package, and projects regularly drop support for older ones as updates roll out. This, much more than raw compute speed, is what ends a card's career in local AI: a RTX 2060 remains physically functional for years after purchase, but some newer tools refuse to install on it because no code was compiled for its generation.

Cards sourced from the QuelLLM hardware database · 20/09/2026
GenerationCardsCompute capacityLocal AI landscape in 2026
PascalGTX 1060, 1070, 1080 Ti6.1Runs tools based on llama.cpp; no FP16 acceleration; abandoned by some recent frameworks
TuringGTX 1660, RTX 2060 at 2080 Ti7.5Minimum required by vLLM (7.0); well supported
AmpereRTX 3050 at 3090 Ti8.6Fully supported; adds BF16
Ada LovelaceRTX 4060 to 40908.9Fully supported; adds FP8 support
BlackwellRTX 5050 to 509012.0Fully supported by recent software; older versions must be updated or recompiled; brings FP4

The last line explains the wave of “my RTX 50 isn't detected” messages: a binary compiled before this architecture existed contains no code for compute capability 12.0. The solution is a newer version of the tool, or PyTorch built for a recent version of CUDA.

!
The « Blackwell » trap: two incompatible families under the same name
NVIDIA sells two different chips under the Blackwell name, and confusing them can break an installation. Consumer GeForce RTX 50 cards use compute capability 12.x (sm_120). Datacenter B100 and B200 cards use compute capability 10.x (sm_100). These are two distinct families: a binary compiled for sm_100 won't run on a RTX 5090, and vice versa. If a fine-tuning tutorial mentions “Blackwell” without specifying sm_120 or sm_100, check before installing a precompiled package: this is a common source of “no kernel image is available” errors on the newest consumer cards.

#CUDA versus ROCm, Metal, and Vulkan

PlatformManufacturerHardwareStatus of local LLMs
CUDANVIDIAGPU NVIDIAThe reference target: it supports everything first
ROCmAMDRecent Radeon and InstinctGood on Linux with supported cards; narrower card list
MetalAppleApple SiliconExcellent with llama.cpp and MLX
VulkanKhronos, open standardAlmost all GPUs, including integrated GPUsWorks almost everywhere with llama.cpp; generally slower than the native solution
SYCL / oneAPIIntelIntel Arc and integrated GPUsFunctional, smallest ecosystem

CUDA's advantage isn't that other platforms are incapable of doing the calculations. It comes from nearly twenty years of libraries, tools, and proven code, built GPU by GPU since 2007 according to the project's official page. For inference with common tools, the gap has narrowed considerably: llama.cpp runs on all five platforms, with similar performance at the same quantization level on comparable hardware. For training, fine-tuning, and research code just published on GitHub, however, CUDA remains the only path that works from day one: it is almost always the first target, and sometimes the only one, for repositories accompanying a new scientific paper, often several months before a port to other platforms appears—if it ever does.

#What this changes when buying

You want everything to work without troubleshooting
Choose a NVIDIA card, and prioritize VRAM capacity over the number of CUDA cores shown on the box: VRAM determines whether a model loads, not raw compute power.
You mainly run GGUFs in Ollama or LM Studio
A Mac with Apple Silicon, or a recent Radeon under Linux with ROCm, is a legitimate choice: llama.cpp supports both platforms properly, and you lose little in day-to-day usability.
You plan to fine-tune, use vLLM, or follow research repositories
Not having CUDA will cost you time every week, between dependencies that fail to compile and notebooks published assuming a GPU NVIDIA is available by default.
You are buying a used card
Avoid anything below compute capability 7.5, regardless of the listed price: beyond being slow, some recent frameworks simply refuse to install.

#FAQ

What does CUDA mean?+
Compute Unified Device Architecture. NVIDIA launched it in 2007 so its GPUs could be programmed with general-purpose code, using languages such as C, C++, or Python, rather than being controlled only through a graphical interface. This shift made the rise of deep learning on graphics cards possible a decade later, since previously every computation had to be disguised as a display operation to run on a GPU.
Do you need to install CUDA to run AI locally?+
Generally, no. Ollama, LM Studio, and comparable applications already bundle the CUDA libraries they need for inference: an up-to-date NVIDIA driver is sufficient, and there is nothing else to download. The full CUDA Toolkit, which takes up several gigabytes, is only needed to compile software from source, such as llama.cpp with GPU acceleration enabled manually.
Does CUDA work on an AMD or Intel card?+
No, never: CUDA is proprietary technology that runs only on NVIDIA hardware, by the manufacturer's commercial choice. The AMD equivalent is called ROCm; for Apple, it is Metal; for Intel, oneAPI; and Vulkan remains the multi-vendor option, supported by most tools based on llama.cpp, including on older hardware or integrated GPUs.
Why don’t nvidia-smi and nvcc show the same version?+
nvidia-smi reports the maximum CUDA version that the installed driver can run, not what is actually installed on the machine. nvcc, on the other hand, reports the version of the CUDA Toolkit actually present, if any. The two numbers are independent, and software compiled for an older CUDA version continues to work without issue with a newer driver.
Are more CUDA cores better for an LLM?+
This mainly speeds up prompt processing and image generation, two tasks that are compute-intensive. An LLM's writing speed, which is what you notice most in practice, depends more on the card's RAM capacity and memory bandwidth: for local language models, VRAM therefore often matters more than the advertised raw CUDA core count.
RTX 50 and Blackwell data center cards: do they have the same computing capability?+
No, and the confusion regularly breaks an installation. Consumer GeForce RTX 50 cards use capability 12.x (sm_120); B100 and B200 datacenter cards use 10.x (sm_100). A binary or precompiled package built for one does not run on the other: a fine-tuning tutorial that merely mentions “Blackwell” without specifying which one should be checked before installing anything.

Recommended hardware: GMKtec EVO-X2 64GB / 1TB (Ryzen AI Max+ 395) — recent NVIDIA card with 16 GB of VRAM, supported by CUDA. All AI hardware →

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.