Beginner 8 minOllama

Ollama Cloud: pricing, reviews, and limitations (2026)

Ollama Cloud offers one simple thing: run models that are too large for your machine—120B, 480B, 671B parameters—on Ollama’s servers, using the same CLI as the local version. It is convenient. But it changes two fundamental things: your prompts leave your computer, and you pay as you go. This guide explains exactly what it is, what it costs, and why self-hosting remains preferable for most use cases.

By Mohamed Meguedmi·Update 2026-09-05·Tested on Windows, macOS, and Linux
i
In brief
Ollama Cloud runs models too large for your machine (120B, 480B, 671B) on Ollama’s servers, using the same CLI as locally. · The offering falls between a monthly plan with an hourly quota and usage-based billing by inference duration for heavy workloads—check pricing at ollama.com/cloud; it changes. · Unlike local use, your prompts leave your machine and pass through servers primarily based in the United States. · For recurring use or sensitive data, self-hosting remains strongly preferable.

#What Ollama Cloud is

Since 2025, Ollama has offered a paid service called Ollama Cloud, which lets you call models hosted on its GPU infrastructure while using exactly the same commands as for a local model. The promise: run massive open-weight models (hundreds of billions of parameters) without needing a Mac Studio Ultra or an H100 cluster.

In practice, you install Ollama as usual, authenticate, and pull a model marked "cloud" (for example, gpt-oss:120b-cloud). Inference no longer runs on your GPU, but on Ollama's servers—your machine only sends the prompt tokens and receives the response tokens.

Example cloud call (general syntax)
# Authentification (une seule fois)
ollama signin

# Lancer un modèle hébergé
ollama run gpt-oss:120b-cloud
i
This is not “local cloud”
Ollama Cloud is not a hybrid mode that runs part of the model locally. It is pure remote inference: your prompts are sent over the Internet, as with the OpenAI or Anthropic API. The only "local" part is the CLI.

#Which models run in the cloud

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The cloud catalog targets precisely what performs poorly locally. It evolves, but the logic remains the same: open-weight or open models, very large models, or multimodal models demanding on VRAM.

100B+ parameter models
Types such as gpt-oss:120b, qwen3-coder:480b, and kimi-k2, which require 60 to 250 GB of VRAM in Q4, are beyond the reach of an individual workstation.
Frontier MoE models
DeepSeek V3/V4, Llama 4 Maverick: the full versions, not the 7B/32B distillations installed locally.
Large specialized models
Coding giants, long-context reasoning (“thinking”) models, high-resolution vision models.

The “normal” models (Mistral 7B, Llama 3.1 8B, Qwen 14B, Phi-4, etc.) are still recommended locally: they fit on a 12 GB RTX 3060 or a MacBook Air, and there is no benefit to offloading them.

#Pricing and limits

Ollama Cloud sits between a flat monthly subscription (with a requests-per-hour quota) and a pay-as-you-go billing model for intensive use. Exact pricing changes over time — check ollama.com/cloud before committing — but a few constants are worth noting.

Hourly quotas
Basic plans limit the number of queries per hour or per day. Comfortable for human chat use, but quickly exhausted by a script looping over 10,000 documents.
Inference-duration billing
Very large models are sometimes billed by GPU time consumed. A long-context request (32k tokens, deep reasoning) mechanically costs more than a “hello.”
No per-token price displayed by default
Unlike OpenAI or Anthropic, the API does not consistently charge per token on the Ollama Cloud side. This makes budgeting less precise.
No consumer-grade SLA
The offering is young, and (at the time of writing) there is no availability guarantee comparable to those of the major clouds.
!
The “almost free” trap
A free quota or inexpensive entry-level plan makes you want to route everything through the cloud. But as soon as you automate things (batch processing, an agent running in a loop, RAG over 100,000 chunks), the bill rises quickly and network latency is added to every call.

#Privacy: what really changes

This is where the difference from self-hosting becomes structural, not incidental. When you use Ollama locally, your prompts never leave port 11434 on localhost — this can be verified at the firewall. In the cloud, they travel over the Internet and are processed on third-party infrastructure.

Your prompts leave the perimeter
Any data sent — a contract excerpt, proprietary source code, patient file, or email — leaves your machine. For a law firm, physician, HR department, or R&D lab, that is a deal-breaker.
Retention policy
Ollama states that it does not train its models on your prompts. However, retention for debugging, abuse detection, or logs varies by plan. Read the terms of service before sending sensitive data.
GDPR and hosting
Ollama Cloud servers are primarily in the United States. For European personal data, this raises questions about transfers outside the EU that self-hosting eliminates in one stroke.
No infrastructure-side audit possible
Locally, you can tcpdump port 11434 and prove that nothing leaves the machine. In the cloud, you have to trust the provider—there is no technical way to verify what is logged.
→
The tcpdump test
For a Ollama self-host, run `sudo tcpdump -i any port 443` while using the model. No outgoing packets should appear. With Ollama Cloud, you’ll see traffic to ollama.com servers. Confidentiality is binary and observable.

#Self-hosted alternatives for each cloud model

For nearly all models offered in the cloud, there is a local alternative that is either equivalent (a smaller model from the same family) or acceptable (another comparable open-weight model). Here are the useful equivalents.

gpt-oss:120b-cloud → llama3.1:8b or qwen2.5:14b locally
For 90% of conversational use cases, an 8B–14B Q4 model on a RTX 3060 12 GB card does the job. The difference is most noticeable on long reasoning tasks or encyclopedic knowledge.
qwen3-coder:480b-cloud → qwen2.5-coder:14b or 32b
The 32B in Q4 (≈19 GB VRAM) runs on a RTX 4090 or a Mac M-Max. For code autocompletion in an IDE, that's more than enough — see the Continue.dev guide.
deepseek-v3:671b-cloud → deepseek-r1:32b or distilled 70b
Distilled versions of DeepSeek R1 reproduce 80–90% of the full model's reasoning on common benchmarks, locally on a 4090 or Mac Studio.
Huge vision models → qwen2.5-vl:7b or llama3.2-vision:11b
For OCR, image tagging, and description, these two models run on 8–12 GB of VRAM. See the guide to local multimodal vision LLMs.
1M-token context → local 128k models + chunking
Very few real-world use cases require 1M tokens at once. A good RAG pipeline with chunking + reranking can process massive corpora with a local 32k model.
i
The quality gap narrows every quarter
A 2026 14B model outperforms a 2024 70B model on most benchmarks. The rationale of "I need a 400B" is becoming increasingly rare. Before paying for the cloud, actually test whether a local 14B is enough—you'll often be surprised.

#Decision table

To decide between Ollama Cloud and self-hosting, ask yourself these five questions in order. The first answer of “self-host” settles the debate.

1. Is the data sensitive?
Proprietary code, customer data, healthcare, legal, HR, R&D → self-host, period.
2. Are you subject to the GDPR or industry-specific regulations?
Hospital, bank, public sector, European data → self-host, or dedicated sovereign cloud (not Ollama Cloud).
3. Is the volume predictable and high?
More than a few thousand calls per day → a €1,500 GPU pays for itself within a few months compared with a variable cloud bill.
4. Do you need to operate offline?
Unreliable connectivity, field demos, air-gapped compliance → self-hosting required.
5. Do you absolutely need a 100B+ model?
If so—and only if so—and after testing a local 32B: Ollama Cloud becomes an option worth considering.
→
Rule of thumb
If you're unsure, start locally with a 14B Q4 model. You'll quickly discover whether its limitations really block your use case—in 80% of cases, they don't. The cloud remains the exception, not the rule.

#When to use one, and when to use the other

#The few good use cases for Ollama Cloud

One-off evaluation of a giant model
You want to test in 10 minutes whether a 480B coder is worth it before investing in a machine — the cloud avoids an impulse purchase.
Client demo with no on-site hardware
A salesperson showing a POC on their laptop without bringing along a Mac Studio.
A single highly demanding task
A 500-page report to summarize once per quarter, with a long context and no confidential data in it.
Learning and exploration
Discover what a 670B model can do, then return to local models knowing what to aim for.

#When self-hosting wins

Any recurring use
Daily coding assistant, internal team chat, RAG over company documents — predictable costs matter.
Any use involving sensitive data
No debate: legal, healthcare, HR, finance, R&D, and proprietary code stay local.
Any automated use
Batch pipelines, looping agents, n8n, ETL scripts: cloud quotas explode and costs become unpredictable.
Any offline or edge use
Raspberry Pi in the workshop, laptop on the road, site without reliable Internet.
Any need for customization
Fine-tuning, a custom Modelfile, team system prompts, deep integration with your tools—the cloud does not allow it.
!
The lock-in trap
Once a team gets used to calling a specific cloud model, switching away is painful: calibrated prompts, output formats, and distinctive behaviors. Investing in self-hosting from the start avoids this dependency—and saves money when prices rise.

#Go further

If you conclude that self-hosting is the right path for your use case (which is likely), these guides will help you get started on the right foot:

Choose the right GPU
The Choosing Your GPU for Local AI guide details the VRAM/budget/performance trade-offs, from the RTX 3060 12 GB to the Mac Studio Ultra.
Understanding quantization
The Choosing Your Quantization guide (Q4, Q5, Q8, FP16) explains how to fit a 32B model on a 16 GB GPU without degrading quality.
Operational privacy
The Privacy Checklist lists the concrete checks needed to ensure that no packet leaves your machine — proof based on facts.
Frequently asked questions about Ollama Cloud
Is Ollama Cloud free?+
There is a free tier, but with restricted usage limits (few simultaneous models, quotas that renew per session and per week). For heavier use, Ollama offers paid tiers through a monthly subscription, providing access to more simultaneous models and quota. Exact pricing changes over time: always check the official ollama.com/pricing page before signing up rather than relying on a fixed figure.
How can you use Ollama without going through the cloud?+
Nothing changes: the standard local installation of Ollama works exactly as before, with no account or subscription. The cloud is an optional feature enabled only when you explicitly choose a model marked "cloud" and authenticate. As long as you use standard models (Llama, Mistral, Qwen, etc.), everything runs locally, with no connection required after the model is downloaded.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.