Intermediate 10 minAPI

Free LLM APIs: the real comparison (and the option locale)

A free LLM API really does exist: several providers offer free access to capable models without a credit card. The catch is what’s in the fine print—tight quotas, throttled throughput, and often your prompts being used to train the next model. This guide takes an honest look at the real offerings, puts a number on what “free” costs in practice, and shows the precise point at which a local Ollama becomes more cost-effective and healthier for a development project.

By Mohamed Meguedmi·Update 2026-09-15·Tested on Windows, macOS, and Linux

#Why this comparison

Looking for a free LLM API is the first instinct when prototyping: you want to test an idea without pulling out your credit card, connect a model to a script, and see whether it holds up. The good news is that free offerings are real and sometimes generous. The bad news is that “free” covers very different realities depending on whether you’re talking about a permanent free tier, an expiring trial credit, or community access with throttled throughput.

The point of this guide is not to tell you “local is better.” It is to give you the figures so you can decide for yourself: what each free offering actually allows, what it requires in return, and at what volume or level of confidentiality a local endpoint becomes the rational choice. Often, the best answer is neither one nor the other, but both, with routing based on the task.

#A tour of free LLM APIs

The Local Copilot Kit

This guide gets you to the model. The kit gets you to the coding copilot in your editor.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

Free offerings can be grouped into four families. Knowing which family an offering belongs to helps avoid unpleasant surprises when the meter hits zero in the middle of a sprint.

Permanent free tier
Free access that does not expire, but capped by requests per minute and per day. This is the case with Google AI Studio (Gemini API) and Groq, which offer open models (Llama, Qwen, gpt-oss) at a free but limited rate.
Aggregators and models :free
OpenRouter exposes dozens of models, some with the “:free” suffix. Access is real, but throughput depends on a shared pool and may drop during peak hours.
Trial credit
An amount offered (often a few dollars) when you create the account, which expires after a few weeks. Useful for a one-off test, but not for a long-running project.
Community inference
Hugging Face Inference and similar services: free access to many models, but with cold starts, queues, and no guaranteed throughput.
i
Exact figures change quickly
The exact thresholds (requests per minute, tokens per day) change every few months for each provider. Always check the official pricing page before sizing a project around them—a free quota can be cut in half overnight.

#Actual quotas, without filtering

The word “free” conceals three separate caps that together determine whether the offering works for your use case. A tier can be generous on one and severely constrained on the other two.

Requests per minute (RPM)
The number of calls allowed per minute. Free tiers often provide around a few dozen RPM—more than enough for a developer testing, but too little to serve multiple users in parallel.
Tokens per day (TPD)
The real structural ceiling. A daily token quota runs out very quickly as soon as you send large prompts, RAG context, or loop over a dataset.
Throughput and latency
On shared free offerings, generation speed is never guaranteed. During off-peak hours, it is smooth; at peak times, latency skyrockets or requests are rejected (429 error).

Concretely: for manual prototyping and a few occasional requests, the free quotas are generous. As soon as you script a batch process—classifying one thousand tickets, summarizing an inbox, generating tests for a repository—you hit the TPD wall within minutes, and throttled throughput turns a 10-minute batch into a one-hour wait punctuated by 429 errors.

!
The production rate-limit trap
A free quota that is enough during development says nothing about production. The day your app has ten simultaneous users, the shared RPM/TPD ceiling becomes the first failure point—and it always hits at the worst time, not during your tests.

#What you are really paying for

A free API is not cost-free: the cost is simply shifted from your wallet to other columns. Three of them weigh heavily on a development project.

Your data
On many free tiers, prompts and responses are retained and may be used to train or improve models. What is acceptable for a test with dummy data is not acceptable with proprietary code, customer data, or personal information (GDPR).
The dependency
Building on a free quota is like building on ground that may give way: terms change, a model is withdrawn, or a tier is removed. Your code, prompts, and settings are calibrated for one provider; migrating costs time you had not planned for.
Unpredictability
Variable latency, queues, outages: it is difficult to promise a quality of service when the core component is outside your control and offers no commitment on the free tier.
!
Read the training clause
Before sending any real data to a free API, look for the wording “we may use your data to improve our models”. On paid tiers, this use is often disabled by default; on the free tier, the opposite is frequently true. If in doubt, assume that anything you send may be read and reused.

#The local option with Ollama

Opposite free-with-conditions, there is truly free: running the model on your own machine. Ollama is the simplest tool for that. It's a daemon that downloads open-weight models and exposes a local HTTP API on http://localhost:11434—including an OpenAI-compatible endpoint, which means code written for a cloud API often works by changing only the base URL.

Terminal — install and run a model
# Installer Ollama (Linux/macOS)
curl -fsSL https://ollama.com/install.sh | sh

# Tirer et lancer un modèle 7-8B quantifié Q4_K_M
ollama pull qwen2.5:7b
ollama run qwen2.5:7b

# Le daemon écoute sur http://localhost:11434

On the code side, the OpenAI-compatible endpoint connects in three lines. No API key to manage, no quota, and no data leaves the machine.

Python — OpenAI client pointed at Ollama
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:11434/v1",  # endpoint local Ollama
    api_key="ollama",  # ignoré en local, mais requis par le client
)

resp = client.chat.completions.create(
    model="qwen2.5:7b",
    messages=[{"role": "user", "content": "Explique la récursion en une phrase."}],
)
print(resp.choices[0].message.content)

The trade-off is hardware. A local model needs VRAM (or unified memory on a Mac Apple Silicon). With Q4_K_M quantization, the memory rule of thumb is easy to remember: a 3B fits in ~2 GB, a 7B in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B requires ~40 GB. A RTX 3060 12 GB runs 7–14B models comfortably; a RTX 4090 24 GB or a Mac M4 Pro with 24–48 GB of unified memory opens the door to 32B models.

Total privacy
No prompt leaves the machine. Proprietary code, customer data, and personal information all stay with you, which radically simplifies GDPR compliance.
No quota
No RPM, no TPD. You can loop over ten thousand documents overnight without tracking usage and without a 429 error.
Zero marginal cost
Once you have the hardware, every request is free. The electricity used by a desktop GPU remains negligible compared with a usage-based API bill.
Stability
No changing conditions, no withdrawn model. The version you downloaded remains identical until you update it.

#The tipping point for going local

The useful question is not “free or local?” but “when does local become the better choice?” Four signals indicate that you have crossed the threshold.

  1. 01
    You process data that you cannot expose
    As soon as proprietary code, customer data, or personal information appears in the prompt, the training clause of a free API becomes a deal-breaker. Local deployment solves the problem at its root: nothing leaves.
  2. 02
    You regularly hit the quotas
    If your scripts end with 429 errors, if you split your batches to stay below the TPD, or if you juggle multiple free accounts, you are already paying in time for what local inference would give back.
  3. 03
    The volume is predictable and sustained
    Regular use—continuous test generation, an internal RAG pipeline, ongoing classification—quickly pays for itself locally. The hardware is a fixed, depreciable cost; usage-based API access is a variable cost that rises with the project's success.
  4. 04
    You want predictable latency
    On a dedicated machine, latency depends only on you, not on the load of a shared service. For an internal tool used all day, that predictability is worth a lot.
→
When free cloud remains the right choice
On-premises isn't always the answer. For a one-off test, access to a very large frontier model your machine can't host, or an occasional and unpredictable workload, a free or pay-as-you-go API is simpler and cheaper than buying a GPU. The best approach is to use both.

#Keep both: smart routing

The healthiest architecture for a development project isn't exclusive: it routes each task to the most suitable endpoint. The local system handles most of the volume and anything sensitive; the cloud is reserved for tasks that genuinely exceed the machine's capabilities.

Going local
High-volume tasks, sensitive data, processing loops, development iterations—anything that must remain confidential. A local 7-14B model covers the vast majority of everyday development needs.
Toward the cloud
Complex reasoning requiring a very large model, a temporary peak workload, or multimodal functionality unavailable locally. Send only what warrants it, and never sensitive data.

Since Ollama exposes an OpenAI-compatible API, this routing is trivial to code: two clients and a selection rule based on the task. To go further, a proxy such as LiteLLM centralizes multiple backends behind a single interface, with automatic fallback from cloud to local (or vice versa) and cost tracking.

Python — local/cloud routing by task
from openai import OpenAI

local = OpenAI(base_url="http://localhost:11434/v1", api_key="ollama")
cloud = OpenAI(base_url="https://api.exemple.com/v1", api_key="VOTRE_CLE")

def router(sensible: bool, gros_raisonnement: bool):
    # Données sensibles OU volume : local par défaut
    if sensible or not gros_raisonnement:
        return local, "qwen2.5:7b"
    # Sinon, cloud pour un modèle plus puissant
    return cloud, "modele-frontiere"

client, model = router(sensible=True, gros_raisonnement=False)
resp = client.chat.completions.create(
    model=model,
    messages=[{"role": "user", "content": "Résume ce ticket interne..."}],
)

#Get started locally in 4 steps

  1. 01
    Install Ollama
    A Linux/macOS command (curl -fsSL https://ollama.com/install.sh | sh) or the official installer on Windows. The daemon starts and listens on http://localhost:11434.
  2. 02
    Choose a model that matches your hardware
    Identify your available VRAM and target a Q4_K_M model that fits with headroom for the context: 7B (~5 GB) for 8–12 GB of VRAM, 14B (~9 GB) for 12–16 GB, 32B (~19 GB) for 24 GB. Below that, a 3B (~2 GB) remains useful for simple tasks.
  3. 03
    Pull and test the model
    ollama pull qwen2.5:7b puis ollama run qwen2.5:7b pour vérifier qu'il répond. Un pull ne se fait qu'une fois ; ensuite le modèle est en cache local.
  4. 04
    Connect your code
    Point your existing OpenAI client at http://localhost:11434/v1. The rest of the code—messages, streaming, function calling—works like a cloud API, with no key or quota.
→
Keep a cloud fallback from the start
Even if you run 100% locally every day, keep the cloud path wired into your code (with a free or pay-as-you-go key). The day you encounter a task that exceeds your machine's capabilities, the fallback is already ready and won't block your progress.

#Go further

Once Ollama is in place, three guides on the site naturally extend this transition. The Python integration of the Ollama REST API covers streaming, JSON mode, and function calling on the local endpoint. The GPU server cost comparison calculates the break-even point between buying hardware and using a cloud API. And for industrial-grade routing between local and cloud, the LiteLLM guide shows how to unify both behind a proxy with fallback and cost tracking.


Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.