Beginner 9 minInterfaces

Cherry Studio: 1 interface for all your LLMs (local +) cloud)

You have Ollama running locally for everyday tasks, an OpenAI credit for complex tasks, and a Mistral key for French. Three different interfaces, three histories, three places to paste your prompts. Cherry Studio is an open-source desktop client that brings all your LLMs—local and cloud—into one app, with custom assistants and MCP support.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Cherry Studio

Cherry Studio (CherryHQ/cherry-studio on GitHub, AGPL-3.0 license) is an Electron desktop app positioned as an LLM provider aggregator. The idea: one interface, one history, your prompts available everywhere—regardless of whether the underlying model is a local Qwen 3.5 on Ollama, GPT-4o at OpenAI, or Claude Sonnet at Anthropic.

For anyone who mixes local and cloud LLMs, Cherry Studio solves a concrete problem: no more juggling LM Studio for local use, ChatGPT.com for OpenAI, Claude.ai for Anthropic, and a Mistral tab for Le Chat. Configure each provider once, then switch between models in the same conversation as needed.

Native multi-provider support
Ollama, LM Studio, OpenAI, Anthropic Claude, Mistral, DeepSeek, Gemini, OpenRouter, Groq, Together, and any OpenAI-compatible endpoint.
Open-source AGPL-3.0
Public code on GitHub. No outbound telemetry by default. You can compile it yourself.
Cross-platform desktop
Windows, macOS (Intel + Apple Silicon), Linux (AppImage, .deb).
Persistent assistants
System prompt + parameters + preconfigured model, reusable like a GPT.
Knowledge bases
File import (PDF, Word, Markdown) with embeddings—integrated RAG with no tinkering.
MCP (Model Context Protocol)
Connect MCP servers to give your LLMs tools (filesystem, web, database).
i
What Cherry Studio is not
This is not an inference engine—it does not download models on its own and does not run LLMs directly. For local use, it needs Ollama or LM Studio as a backend. It is a client, not a server.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
64-bit system
Windows 10/11, macOS 11+, recent Linux. No official mobile version (Cherry Studio is desktop-only).
4 GB of free RAM
The app itself is lightweight (Electron), but if you also run Ollama with a loaded 7B model, plan on 8 GB total.
At least one LLM backend
Either a Ollama running on localhost:11434 or an API key from a cloud provider (OpenAI, Anthropic, Mistral...). Without that, the app is empty.
(Optional) Node.js for MCP
Several MCP servers install via npx — Node 18+ recommended if you want to use this feature.

#1. Windows / macOS / Linux installation

The project publishes signed binaries on GitHub’s Releases page. Get the latest stable version directly from the official repository.

GitHub releases
https://github.com/CherryHQ/cherry-studio/releases
  1. 01
    Windows
    Download Cherry-Studio-x.y.z-x64-setup.exe (NSIS installer) or the portable version. Double-click it and follow the wizard. No external dependencies need to be installed.
  2. 02
    macOS
    Download the .dmg that matches your architecture: Cherry-Studio-x.y.z-arm64.dmg for Apple Silicon (M1/M2/M3/M4), -x64.dmg for Intel. Drag the app into /Applications. On first launch, right-click → Open to bypass Gatekeeper if the app is not notarized.
  3. 03
    Linux
    Download the .AppImage (universal) or the .deb (Debian/Ubuntu). For the AppImage: chmod +x Cherry-Studio-x.y.z.AppImage then double-click. For the .deb: sudo dpkg -i Cherry-Studio-x.y.z.deb.
→
Verify binary integrity
On the Releases page, each archive has a published SHA-256 hash. On Linux: sha256sum Cherry-Studio-*.AppImage and compare them. On macOS: shasum -a 256 Cherry-Studio-*.dmg. It takes 10 seconds and protects you from a compromised release.

#2. Connect Ollama locally

On first launch, Cherry Studio opens to a blank screen—no model configured. Start with the simplest option: point it to your local Ollama.

First, make sure Ollama is running and listening on its standard port.

Check Ollama
curl http://localhost:11434/api/tags

You should receive a JSON response with the list of installed models. If curl returns a connection error, start the Ollama service before continuing.

  1. 01
    Open settings
    In Cherry Studio, click the Settings icon (bottom left), then Model Providers.
  2. 02
    Enable Ollama
    Find Ollama in the list of providers and switch the toggle to ON. The API Host field is filled in by default with http://localhost:11434 — don't touch anything if Ollama is running on the same machine.
  3. 03
    Synchronize the models
    Click Manage on the right, then click the refresh button. Cherry Studio queries Ollama and lists all your installed models (qwen3.5, granite4.2, gemma4...). Check the ones you want to appear in the selector.
  4. 04
    Test
    Return to the Chat view, create a new conversation, choose a Ollama model from the menu at the top, and send a message. The response should stream in the window.
i
Ollama on another machine
If your Ollama runs on a server on the local network, replace localhost with its IP (e.g., http://192.168.1.42:11434). On the server side, Ollama must be launched with OLLAMA_HOST=0.0.0.0:11434 to listen on all interfaces; otherwise, it will refuse remote connections.

#3. Add cloud providers alongside them

Cherry Studio becomes useful when you mix and match. You can have 5 active providers simultaneously and switch between them with one click in the model selector.

OpenAI
API key on platform.openai.com. Cherry Studio exposes GPT-4o, GPT-4 Turbo, o1, etc.
Anthropic Claude
API key at console.anthropic.com. Claude Sonnet 4.6, Haiku, and Opus models available.
Mistral
API key on console.mistral.ai. Access to Mistral Large, Codestral, and Pixtral.
DeepSeek
API key at platform.deepseek.com. Access to DeepSeek V3, R1 (reasoning).
OpenRouter
A single key to access 200+ models through a unified proxy—convenient for quick testing.
Custom endpoint
For any OpenAI-compatible server: vLLM, LM Studio in server mode, Together, Fireworks, llama.cpp server.
  1. 01
    Settings → Model Providers
    Enable every provider you want to use.
  2. 02
    Paste the API Key
    API Key field. Cherry Studio encrypts keys in the app’s local storage (not in plain text in a visible config file).
  3. 03
    Test the connection
    The Check button next to the key sends a test request. Green = OK, red = invalid key or blocked network.
  4. 04
    Choose which models to expose
    For each provider, check the models you want to see in the selector. Uncheck the ones you never use so they don't overwhelm the menu.
!
Your API keys stay local
Cherry Studio is a client: your keys are stored on your machine and used to communicate directly with provider APIs. But that also makes them your responsibility — losing your laptop means losing access to your keys. If you work with others, create separate keys for each workstation instead of sharing them.

#4. Custom assistants

A Cherry Studio assistant = a system prompt + a default model + settings (temperature, context) + optionally an attached knowledge base. It is the local equivalent of OpenAI's GPTs, except that you remain in control.

  1. 01
    Click + in the Assistants sidebar
    A creation panel opens. Name, emoji, short description.
  2. 02
    Define the system prompt
    This is the assistant’s core: its role, constraints, and expected tone. Example: “You are a French essay grader. You respond by pointing out errors one by one and providing the correction. No paraphrasing, no sugarcoating.”
  3. 03
    Choose the default model
    You can select any model from any configured provider. For French text correction: mistral-large via API, or a local mistral-small through Ollama for privacy.
  4. 04
    Adjust temperature and context
    0.2 for a deterministic task (correction, extraction), 0.7–1.0 for creative writing. The context window depends on the model (up to 128k for newer models).
  5. 05
    Save
    The assistant appears in the sidebar. Click it to start a new conversation preconfigured with this setup.
→
Share an assistant
Cherry Studio exports assistants as JSON. Useful for versioning your prompts in a git repository or sharing them with a team without manually copying the system prompt and parameters.

#5. MCP (Model Context Protocol) support

MCP is the protocol published by Anthropic in late 2024 to standardize how an LLM accesses external tools: reading a file, querying a database, or calling an API. Cherry Studio is one of the first mainstream desktop clients to support it natively.

In practice, you declare an MCP server in Cherry Studio, and all your tool-use-compatible models (Claude, GPT-4, some Qwen 3.5 models, and Granite 4.2) can call it during a conversation.

Example: MCP filesystem in Settings → MCP Servers
{
  "filesystem": {
    "command": "npx",
    "args": [
      "-y",
      "@modelcontextprotocol/server-filesystem",
      "/home/user/documents"
    ]
  }
}

After it is added, the assistant can list, read, and write files in the authorized folder. Ask it, "List the .md files in my documents folder and summarize the most recent one"—it invokes the tool automatically.

!
MCP = pay attention to what you authorize
An MCP filesystem server pointed at /home/user gives the LLM read/write access to your entire home directory. Always limit the scope to what is strictly necessary. The same applies to MCP servers exposing a database or shell—start in read-only mode.

#Cherry Studio vs. Chatbox vs. Jan

All three are open-source desktop clients for multi-provider LLMs. The differences:

Cherry Studio
The most comprehensive aggregation layer: 15+ cloud providers, Ollama/LM Studio locally, assistants, knowledge bases, MCP. Dense interface, with a steeper learning curve.
Chatbox
More minimalist and accessible. Fewer providers, no MCP, and no equally advanced knowledge bases. Great if you just want a clean multi-API client without the overhead.
Jan
More local-first in its approach: Jan includes its own inference engine (based on llama.cpp), so it can run models without depending on Ollama. Fewer cloud providers than Cherry, and no MCP at the time of writing.

In summary: Cherry Studio if you mix a lot of local and cloud models and want MCP. Jan if you want an all-in-one local solution without Ollama. Chatbox if you want the simplest client.


#Troubleshooting

No Ollama model is displayed
Check that curl http://localhost:11434/api/tags responds, then click Manage → the refresh button in provider Ollama. If Ollama is running but not responding, restart the service.
401 error on OpenAI/Anthropic
Invalid or expired key. Regenerate it from the provider’s console and paste it again. Also check that your account has credit.
Connection error on the cloud behind a corporate proxy
Cherry Studio does not consistently read the HTTPS_PROXY environment variables. Configure the proxy under Settings → General → Network.
MCP server does not start
Node.js is missing or too old. Install Node 18+ (nvm recommended), then relaunch Cherry Studio. MCP logs are available via View → Toggle Developer Tools → Console.
Slow / laggy app
Too much accumulated history. Long conversations weigh down the local index. Archive or delete old ones via Settings → Data Management.
Response cut off midway
The model hit its output token limit. Increase Max Tokens in the assistant settings (4096 or more for long responses).

#Go further

Cherry Studio is mainly interesting as a unified frontend. To dig into the backend running underneath:

If you don't already have Ollama
“What is Ollama and how does it work?” explains the daemon that Cherry Studio queries locally—useful for understanding why a model fails to load.
If you compare Cherry with other interfaces
“LM Studio vs Ollama: which should you choose in 2026?” and the guide “Msty: the elegant local LLM interface” cover the direct alternatives in the unified-interface space.
If you want to share resources further
“LiteLLM: a unified local and cloud proxy” goes further than Cherry Studio by operating server-side, with automatic fallback between a local Ollama and a cloud API.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.