Beginner 14 minFree

Best free local AI: 2026 roundup, even without GPU

Direct response

Yes: Ollama, LM Studio, Jan, and GPT4All can be downloaded for free, and Qwen 3.5, Gemma 4, Granite 4.2, and gpt-oss are released under the Apache 2.0 license. The right choice depends on memory: a 2 to 4 GB file suits an 8 GB machine, 6 to 8 GB requires 16 GB, and above 14 GB you need 24 GB of VRAM or 32 GB of memory.

A free local AI comes down to two downloads: software (Ollama, LM Studio, Jan) and an open-weight language model. The real selection criterion is not the model's name but its file size compared with your memory. This page provides the actual sizes published by Ollama, a speed calculation for machines without a graphics card, and the cases where a so-called “local” AI still sends your data elsewhere.

Choosing a machine? Our picks by budget →

By Mohamed Meguedmi·Update 2026-09-29·Tested on Windows, macOS, and Linux

#Get started for free based on your machine

There is no single best free local AI: the right choice is the largest model whose file fits in your memory with headroom, launched from Ollama or LM Studio. On an 8 GB machine without a graphics card, target a 2 to 4 GB file, such as Granite 4.2 3B (2.2 GB) or Qwen 3.5 4B (3.4 GB). With 16 GB, Qwen 3.5 9B (6.6 GB), Granite 4.2 8B (5.3 GB), or Gemma 4 12B (7.6 GB) are reasonable starting points. Starting at 24 GB of VRAM or 32 GB of memory, gpt-oss 20B (14 GB) and models with 27 billion parameters (18 GB) become accessible. These sizes are those of the files announced by Ollama, not the billions of parameters in the name: it is the file, plus context and system headroom, that must fit. These are starting points, not guarantees of quality or speed: try the model on your own tasks.

Models to try based on memory (file sizes published by Ollama, last checked September 29, 2026)
Machine memoryModels to try (tag Ollama)File sizeAdvertised context
8 GB, no GPUgranite4.2:3b ; qwen3.5:4b2.2 GB; 3.4 GB128 k; 256 k tokens
16 GBgranite4.2:8b; qwen3.5:9b; gemma4:12b5.3 GB; 6.6 GB; 7.6 GB128k; 256k; 256k tokens
24 GB of VRAM or 32 GB of memorygpt-oss:20b ; qwen3.6:27b ; qwen3.8:27b ; gemma4:26b14 GB; 18 GB; 18 GB; 19 GB128 k; 256 k; 256 k; 256 k tokens
32 GB and up, for codingqwen3-coder:30b19 GB256 k tokens

The stated context is a maximum, not what will be allocated. Ollama starts with 4k tokens when the graphics card has less than 24 GiB of VRAM, and recommends at least 64,000 tokens for web search, agents, and code. Each increase consumes additional memory beyond the file itself.

#What is local AI?

The Local AI Kit

You know which free AI runs on your machine. The Local AI Kit takes you all the way to an assistant you use every day: guided installation in one hour (ch. 1), the right model for your exact machine (ch. 3), and what really works without a graphics card (ch. 12).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

A local AI is a language model (an LLM, such as Qwen, Gemma, or Granite) whose files are on your drive and that software installed on your machine runs. Computation happens on your processor or graphics card, not on a vendor’s servers. LM Studio states this in its documentation: what you type while chatting with a model never leaves the device.

The models concerned are described as open-weight models: their parameters can be downloaded. This is not synonymous with open source, which depends on the license. Google points this out for Gemma: Gemma 4, released on April 2, 2026, is the first in the family under an Apache 2.0 license approved by the OSI; previous generations were not under this license. Before reusing a model in a product, read its license.

i
In two words
Local AI = a model downloaded to your machine plus software that runs it. Once the model is on disk, chatting does not require a connection.

The opposite is an online service such as ChatGPT, Gemini, or Claude: your messages are sent to the provider's servers. Between the two are hybrid cases, detailed below.

#Why switch to local AI?

Privacy
Your documents and prompts stay on the machine as long as you enable neither a hosted model nor web search. This is useful for sensitive, legal, or medical work.
Offline
The downloaded model answers without a connection, on a train, or on an isolated network.
No message quota
No plan caps your conversations. The only limit is the speed of your hardware.
Control
You choose the model, its quantization, and the context, and replace it when a better one comes out.

The tradeoff is real: a local model has to fit in your memory, so it is smaller than models from online services.

#Free local AI: really free?

Yes, for the software and models mentioned here: no subscription is needed to chat with a model installed on your machine. Three nuances help avoid surprises: software licenses differ (LM Studio is free but not open source), the cost shifts to storage and hardware, and some vendors also sell online offerings.

The disk
On Windows, Ollama requires at least 4 GB for installation, plus space for the models, which its documentation says can reach from several dozen to several hundred GB.
The hardware
A 14 GB file makes no sense with 8 GB of memory: test small before you buy.
Hosted offerings
Ollama's free plan includes local execution, as well as hosted models billed by usage, with paid plans above that. This is not local AI.
i
What “free” means here
Software without subscriptions, models without licensing fees, and no message quotas. You are responsible for the disk, electricity, and memory.

#Top free tools

Four free tools compared (sources: official repositories and documentation)
ToolLicenseInterfaceSupported systems
OllamaMITTerminal, local API, and desktop applicationWindows 10 22H2 or later; macOS 14 or later (Apple Silicon, or Intel in CPU-only mode); Linux
LM StudioProprietary, freeFull graphical interfacemacOS 14+ on Apple Silicon only; Windows x64 and ARM; Linux
JanApache 2.0Desktop applicationWindows, macOS, Linux
GPT4AllMITDesktop applicationWindows, macOS 12.6+, Linux x86-64
Ollama
The developers’ choice: launch a model with one command, with a local API on port 11434.
LM Studio
Model catalog, one-click downloads, and a chat window: the simplest way to get started. Free at home and at work since July 8, 2025, but proprietary: its terms grant a usage license, not open source code.
Jan
Downloads models from Hugging Face and opens a local server. Its repository states that it can also connect to OpenAI, Anthropic, or others: it's up to you to stay local.
GPT4All
Its repository advertises “no API calls or GPUs required.” Its documentation targets models with 3 to 13 billion parameters and specifies 8 GB of RAM for Llama 3 8B. No release has come out since 3.10.0 in February 2025: verify that it supports the model you want.

#Without a graphics card: how the CPU performs

Yes, a model can run entirely on the processor and system memory. Ollama confirms this in its FAQ: the Processor column in ollama ps displays “100% CPU” when the model is loaded entirely into system memory. The real question is speed, not feasibility.

#Why the processor is slow: a calculation

Each generated token forces the machine to reread most of the model from memory: speed is limited more by data access speed than by compute power. A theoretical ceiling is obtained by dividing memory bandwidth (GB/s) by file size (GB). The table uses an example value of 50 GB/s: this is an illustrative calculation, not a measurement.

Theoretical generation ceiling for memory at 50 GB/s (calculation example, not a measurement)
ModelFile sizeTheoretical ceiling
granite4.2:3b2.2 GBapproximately 23 tokens per second
granite4.2:8b5.3 GBabout 9 tokens per second
qwen3.5:9b6.6 GBapproximately 8 tokens per second

Reality falls short. The TensorFoundry blog reports that on an M2 Max Mac at 400 GB/s, the calculated ceiling is about 87 tokens per second for an 8B in 4-bit, with 64.9 measured, or 75%. Practical consequence: without a GPU, file size is the primary speed lever, since the ceiling depends directly on it. Mixture-of-experts (MoE) models change the equation because they activate only part of their parameters per token: Qwen3-Coder 30B-A3B activates 3.3 billion out of 30. Worth testing if memory is sufficient for the entire file.

#Depending on your system

Windows
Ollama requires Windows 10 22H2 or later. LM Studio requires an AVX2 processor and recommends 16 GB of RAM and 4 GB of VRAM.
Mac Apple Silicon
The processor and GPU share unified memory. LM Studio recommends 16 GB or more; on 8 GB, it recommends smaller models and modest contexts.
Intel Mac
LM Studio does not support it. Ollama runs on x86 but CPU-only: plan on small models.
Linux
Ollama and LM Studio are available; the latter is distributed as an AppImage and requires Ubuntu 20.04 or later.

#The best free models of 2026

This selection is based on what Ollama publishes and what the editors document, not on an in-house test.

Qwen 3.5 (Alibaba)
Multimodal family (text and image) ranging from 0.8 to 122 billion parameters, with a 256k-token context and an Apache 2.0 license. Alibaba claims 201 languages and dialects. The 4B (3.4 GB) and 9B (6.6 GB) versions target small machines.
Qwen 3.6 and 3.8 (Alibaba)
Qwen 3.6 is available in 27B (18 GB) and 35B (23 GB), while Qwen 3.8 is available in 27B (18 GB): two families focused on coding and agents. Qwen3.8's reasoning mode is enabled by default and can be disabled per request: it adds tokens before the response.
Granite 4.2 (IBM)
3B (2.2 GB), 8B (5.3 GB), and 30B versions, with a 128k-token context and an Apache 2.0 license. French is one of the twelve supported languages: that is a signal, not a quality measure. An integrated thinking mode can be disabled with the /set nothink command in Ollama.
Gemma 4 (Google)
Tailles E2B, E4B, 12B, 26B (à experts, 3,8 milliards de paramètres actifs) et 31B, sous Apache 2.0 depuis avril 2026. Le « E » signifie « effective » : gemma4:e2b compte 2,3 milliards de paramètres effectifs mais 5,1 avec les plongements, et son fichier pèse 7,2 Go chez Ollama. Un nom en 2B ne garantit pas un petit fichier.
gpt-oss 20B (OpenAI)
Open-weight model under the Apache 2.0 license, quantized in MXFP4 by OpenAI: Ollama indicates that it runs with 16 GB of memory (14 GB file). 128k-token context.
Qwen3-Coder 30B-A3B (Alibaba)
Code specialist: 30 billion parameters overall, 3.3 activated, 256 k-token context, 19 GB file.

#How to read a benchmark correctly

Model cards publish scores, but those are the publisher's scores, using its own protocol. The Gemma 4 card gives 69.4% on MMLU Pro for the E4B version versus 85.2% for the 31B: size matters, and a score says nothing about your task. Another vocabulary trap: “free LLM” can also refer to quota-based online services, where your data is sent to the provider.

#Choose based on usage: chat, code, agents

Chat, summarize, and write in French
Granite 4.2 8B, Qwen 3.5 9B, or Gemma 4 12B on 16 GB. Compare them on three of your own texts: no spec sheet replaces this test.
Coder
Qwen3-Coder 30B-A3B is presented by Alibaba as its most agent-oriented coding model. The constraint is memory: coding tools require a long context, which adds to the 19 GB file and leaves little room on 24 GB.
Agents and tools
Ollama can launch agents with a local model, such as Claude Code with a Qwen model. This runs the Claude Code tool, not a Claude model: the model that responds is still the one you downloaded.
No restrictions
The original models are trained to refuse certain requests. So-called abliterated community variants remove this capability: a Hugging Face post explains that refusal is carried by a specific direction in the activations that can be neutralized. Quality varies by variant, and the law applies to your use cases.

#Install free AI, step by step

  1. 01
    Check free memory
    Close the browser and resource-intensive applications: available memory, not total capacity, determines the model.
  2. 02
    Install the software
    Ollama if you're comfortable with a command, LM Studio for a graphical interface. Choose your system version: Windows, macOS, or Linux.
  3. 03
    Download a model of the right size
    Refer to the table in the first section: 8 GB of memory, a 2 to 4 GB file; 16 GB, a 5 to 8 GB file. With 8 GB, a 6.6 GB file leaves no margin.
  4. 04
    Launch, then monitor
    Chat, then run ollama ps in a second terminal to see whether the model is running on the GPU, CPU, or both.
Examples: a lightweight model, then a 16 GB model
# Machine de 8 Go
ollama run qwen3.5:4b

# Machine de 16 Go
ollama run qwen3.5:9b

# Vérifier le placement CPU ou GPU et le contexte alloué
ollama ps
Agents and code: increase the context (uses more memory)
OLLAMA_CONTEXT_LENGTH=64000 ollama serve

#What can leave the machine despite the word “local”

“Local” describes where the model runs, not what the rest of the software does. Four situations can send out or expose data.

Ollama-hosted models
The application lets you use models running in the Ollama cloud, which then processes your prompts and responses; Ollama sees nothing you run locally. To disable it: variable OLLAMA_NO_CLOUD=1 or option disable_ollama_cloud.
Jan's connectors
Jan can connect to OpenAI, Anthropic, Mistral, or Groq. With a downloaded model, execution remains local; with a connector, the conversation is sent to the selected provider.
Downloading the models
LM Studio sends network requests to find and download models and check for updates. The conversation, your documents, and the local server stay on the device.
Network exposure
Ollama listens on 127.0.0.1 by default: only your computer can access it. The OLLAMA_HOST variable changes the address; expose it only if you know who can connect.
!
Identify a hosted model
In Ollama, a tag ending in cloud, such as gemma4:31b-cloud, does not run locally. If privacy is your reason for running locally, disable these features in the configuration.
Frequently asked questions
What is the best free local AI for beginners?+
Ollama or LM Studio with a model whose file fits comfortably in your memory: Qwen 3.5 4B (3.4 GB) on 8 GB, Qwen 3.5 9B (6.6 GB) or Granite 4.2 8B (5.3 GB) on 16 GB. Compare two models on your own texts.
Can you get a free local AI without a GPU?+
Yes. A model runs on the processor and RAM, and Ollama then shows 100% CPU. The limitation is speed, which depends on memory bandwidth divided by file size. Choose a 2 to 4 GB model and keep the context short.
Are LM Studio and Ollama really free?+
For local use, yes. LM Studio has been free at home and at work since July 8, 2025, but it is not open source. Ollama is under the MIT license; its hosted models are usage-based, with paid plans, and fall outside the scope of free local use.
Does my local AI send data over the Internet?+
Not when the conversation is with a model loaded on your machine: LM Studio indicates that what you enter does not leave the device. Data does leave with a hosted model, an online connector, or a cloud tag. Disable these options if you want strictly local use.
Can you run Claude or ChatGPT locally?+
The Claude Code tool can be launched with a local model, such as Qwen, via Ollama, but the model that responds is then the one you downloaded. OpenAI publishes gpt-oss, with open weights under the Apache 2.0 license: it's the closest path to ChatGPT locally.
Is there an unrestricted local AI?+
So-called abliterated community variants remove a model's refusal capability and can be found on Hugging Face. Their quality varies from one variant to another, and what you do with them remains subject to the law. The dedicated guide explains what they change and the precautions to take.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.