Free LLM APIs: the real comparison (and the option locale)
A free LLM API really does exist: several providers offer free access to capable models without a credit card. The catch is what’s in the fine print—tight quotas, throttled throughput, and often your prompts being used to train the next model. This guide takes an honest look at the real offerings, puts a number on what “free” costs in practice, and shows the precise point at which a local Ollama becomes more cost-effective and healthier for a development project.
#Why this comparison
Looking for a free LLM API is the first instinct when prototyping: you want to test an idea without pulling out your credit card, connect a model to a script, and see whether it holds up. The good news is that free offerings are real and sometimes generous. The bad news is that “free” covers very different realities depending on whether you’re talking about a permanent free tier, an expiring trial credit, or community access with throttled throughput.
The point of this guide is not to tell you “local is better.” It is to give you the figures so you can decide for yourself: what each free offering actually allows, what it requires in return, and at what volume or level of confidentiality a local endpoint becomes the rational choice. Often, the best answer is neither one nor the other, but both, with routing based on the task.
#A tour of free LLM APIs
This guide gets you to the model. The kit gets you to the coding copilot in your editor.
- Lifetime online access
- PDF + files
- Lifetime updates
Free offerings can be grouped into four families. Knowing which family an offering belongs to helps avoid unpleasant surprises when the meter hits zero in the middle of a sprint.
- Permanent free tier
- Free access that does not expire, but capped by requests per minute and per day. This is the case with Google AI Studio (Gemini API) and Groq, which offer open models (Llama, Qwen, gpt-oss) at a free but limited rate.
- Aggregators and models :free
- OpenRouter exposes dozens of models, some with the “:free” suffix. Access is real, but throughput depends on a shared pool and may drop during peak hours.
- Trial credit
- An amount offered (often a few dollars) when you create the account, which expires after a few weeks. Useful for a one-off test, but not for a long-running project.
- Community inference
- Hugging Face Inference and similar services: free access to many models, but with cold starts, queues, and no guaranteed throughput.
#Actual quotas, without filtering
The word “free” conceals three separate caps that together determine whether the offering works for your use case. A tier can be generous on one and severely constrained on the other two.
- Requests per minute (RPM)
- The number of calls allowed per minute. Free tiers often provide around a few dozen RPM—more than enough for a developer testing, but too little to serve multiple users in parallel.
- Tokens per day (TPD)
- The real structural ceiling. A daily token quota runs out very quickly as soon as you send large prompts, RAG context, or loop over a dataset.
- Throughput and latency
- On shared free offerings, generation speed is never guaranteed. During off-peak hours, it is smooth; at peak times, latency skyrockets or requests are rejected (429 error).
Concretely: for manual prototyping and a few occasional requests, the free quotas are generous. As soon as you script a batch process—classifying one thousand tickets, summarizing an inbox, generating tests for a repository—you hit the TPD wall within minutes, and throttled throughput turns a 10-minute batch into a one-hour wait punctuated by 429 errors.
#What you are really paying for
A free API is not cost-free: the cost is simply shifted from your wallet to other columns. Three of them weigh heavily on a development project.
- Your data
- On many free tiers, prompts and responses are retained and may be used to train or improve models. What is acceptable for a test with dummy data is not acceptable with proprietary code, customer data, or personal information (GDPR).
- The dependency
- Building on a free quota is like building on ground that may give way: terms change, a model is withdrawn, or a tier is removed. Your code, prompts, and settings are calibrated for one provider; migrating costs time you had not planned for.
- Unpredictability
- Variable latency, queues, outages: it is difficult to promise a quality of service when the core component is outside your control and offers no commitment on the free tier.
#The local option with Ollama
Opposite free-with-conditions, there is truly free: running the model on your own machine. Ollama is the simplest tool for that. It's a daemon that downloads open-weight models and exposes a local HTTP API on http://localhost:11434—including an OpenAI-compatible endpoint, which means code written for a cloud API often works by changing only the base URL.
On the code side, the OpenAI-compatible endpoint connects in three lines. No API key to manage, no quota, and no data leaves the machine.
The trade-off is hardware. A local model needs VRAM (or unified memory on a Mac Apple Silicon). With Q4_K_M quantization, the memory rule of thumb is easy to remember: a 3B fits in ~2 GB, a 7B in ~5 GB, a 14B in ~9 GB, a 32B in ~19 GB, and a 70B requires ~40 GB. A RTX 3060 12 GB runs 7–14B models comfortably; a RTX 4090 24 GB or a Mac M4 Pro with 24–48 GB of unified memory opens the door to 32B models.
- Total privacy
- No prompt leaves the machine. Proprietary code, customer data, and personal information all stay with you, which radically simplifies GDPR compliance.
- No quota
- No RPM, no TPD. You can loop over ten thousand documents overnight without tracking usage and without a 429 error.
- Zero marginal cost
- Once you have the hardware, every request is free. The electricity used by a desktop GPU remains negligible compared with a usage-based API bill.
- Stability
- No changing conditions, no withdrawn model. The version you downloaded remains identical until you update it.
#The tipping point for going local
The useful question is not “free or local?” but “when does local become the better choice?” Four signals indicate that you have crossed the threshold.
- 01You process data that you cannot exposeAs soon as proprietary code, customer data, or personal information appears in the prompt, the training clause of a free API becomes a deal-breaker. Local deployment solves the problem at its root: nothing leaves.
- 02You regularly hit the quotasIf your scripts end with 429 errors, if you split your batches to stay below the TPD, or if you juggle multiple free accounts, you are already paying in time for what local inference would give back.
- 03The volume is predictable and sustainedRegular use—continuous test generation, an internal RAG pipeline, ongoing classification—quickly pays for itself locally. The hardware is a fixed, depreciable cost; usage-based API access is a variable cost that rises with the project's success.
- 04You want predictable latencyOn a dedicated machine, latency depends only on you, not on the load of a shared service. For an internal tool used all day, that predictability is worth a lot.
#Keep both: smart routing
The healthiest architecture for a development project isn't exclusive: it routes each task to the most suitable endpoint. The local system handles most of the volume and anything sensitive; the cloud is reserved for tasks that genuinely exceed the machine's capabilities.
- Going local
- High-volume tasks, sensitive data, processing loops, development iterations—anything that must remain confidential. A local 7-14B model covers the vast majority of everyday development needs.
- Toward the cloud
- Complex reasoning requiring a very large model, a temporary peak workload, or multimodal functionality unavailable locally. Send only what warrants it, and never sensitive data.
Since Ollama exposes an OpenAI-compatible API, this routing is trivial to code: two clients and a selection rule based on the task. To go further, a proxy such as LiteLLM centralizes multiple backends behind a single interface, with automatic fallback from cloud to local (or vice versa) and cost tracking.
#Get started locally in 4 steps
- 01Install OllamaA Linux/macOS command (curl -fsSL https://ollama.com/install.sh | sh) or the official installer on Windows. The daemon starts and listens on http://localhost:11434.
- 02Choose a model that matches your hardwareIdentify your available VRAM and target a Q4_K_M model that fits with headroom for the context: 7B (~5 GB) for 8–12 GB of VRAM, 14B (~9 GB) for 12–16 GB, 32B (~19 GB) for 24 GB. Below that, a 3B (~2 GB) remains useful for simple tasks.
- 03Pull and test the modelollama pull qwen2.5:7b puis ollama run qwen2.5:7b pour vérifier qu'il répond. Un pull ne se fait qu'une fois ; ensuite le modèle est en cache local.
- 04Connect your codePoint your existing OpenAI client at http://localhost:11434/v1. The rest of the code—messages, streaming, function calling—works like a cloud API, with no key or quota.
#Go further
Once Ollama is in place, three guides on the site naturally extend this transition. The Python integration of the Ollama REST API covers streaming, JSON mode, and function calling on the local endpoint. The GPU server cost comparison calculates the break-even point between buying hardware and using a cloud API. And for industrial-grade routing between local and cloud, the LiteLLM guide shows how to unify both behind a proxy with fallback and cost tracking.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.