Beginner 11 minInterfaces

Msty: the elegant local LLM interface (Mac, Windows, Linux)

Msty is a desktop app focused entirely on the user experience for running LLMs locally. Where Open WebUI requires Docker and LM Studio focuses on the engine, Msty offers everything in a single installer: a local model, cloud models (OpenAI, Claude, Gemini) side by side, native RAG through Knowledge Stacks, workspaces, and split chat for comparing two models. This guide shows how to install and use Msty on Mac, Windows, and Linux for serious local LLM use.

By Mohamed Meguedmi·Update 2026-08-27·Tested on Windows, macOS, and Linux

#Why Msty

Msty (msty.app) targets a specific profile: someone who wants a clean graphical app, without Docker or a CLI, but with real features beyond simple chat. Three differentiators sum up the tool.

Everything in a single app
Built-in inference engine (based on llama.cpp), client for Ollama if you already installed it, and cloud connectors—no stack to assemble.
Knowledge Stacks
RAG is built in and requires no code. Drag and drop PDFs, Markdown, entire folders, or even URLs; Msty handles chunking, embeddings, and retrieval.
Workspaces and Split Chat
Several isolated workspaces (personal, work, R&D), with the ability to ask the same question to 2-4 models in parallel in the same view.
i
Budget model
Msty is freemium. The free version covers the vast majority of use cases (local chat, cloud, basic Knowledge Stacks, split chat with 2 models). The paid Aurum version unlocks unlimited split chat, an advanced prompts library, and a few pro options. According to the publisher’s policy, neither version collects telemetry from your conversations.

#Prerequisites

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
OS
macOS 11+ (Apple Silicon or Intel), Windows 10/11 (x64), Linux x64 (AppImage and .deb/.rpm packages).
Minimum RAM
8 GB for a 3B Q4_K_M model, 16 GB for 7B, 32 GB for a comfortable 14B, 64 GB for targeting 32B.
GPU
Optional but recommended. NVIDIA with CUDA, AMD with ROCm/Vulkan, Apple Silicon with Metal — Msty detects and uses them automatically.
Storage
Allow 10-50 GB depending on the number of models. GGUF Q4 files range from 2 GB (3B) to 40 GB (70B).
→
Q4 VRAM by model size
With Q4_K_M quantization (Msty's default): 3B≈2 GB · 7B≈5 GB · 14B≈9 GB · 32B≈19 GB · 70B≈40 GB. Beyond your VRAM capacity, Msty offloads to system RAM—usable but slow (dropping from 50 to 5 tokens/sec).

#1. Installation

Msty can only be downloaded from the official website. There is no package manager support yet (neither Homebrew nor winget), which makes it easier to verify the source.

Official website
https://msty.app
!
Download only from msty.app
The publisher (Sapling Inc.) does not sign through the official stores. Check the URL carefully—third-party sites sometimes redistribute modified versions. The official site is msty.app, not msty.io or other variants.
  1. 01
    Download the installer
    Go to msty.app and choose your OS. Three variants are available: Online (reduced, cloud-only), Local (recommended, includes the inference engine), and Local + GPU (with CUDA drivers bundled for Windows).
  2. 02
    Install (Mac)
    Open the .dmg, drag Msty to Applications, then launch it for the first time with right-click → Open (Gatekeeper). macOS asks for confirmation because the app is not notarized by the App Store.
  3. 03
    Install (Windows)
    Run the .exe. Choose the GPU variant if you have a recent NVIDIA—it includes the CUDA runtimes and avoids manually wrestling with the drivers. Otherwise, the Local CPU/Vulkan version is sufficient.
  4. 04
    Install (Linux)
    AppImage: chmod +x Msty-*.AppImage, then ./Msty-*.AppImage. .deb: sudo dpkg -i msty_*.deb. .rpm: sudo rpm -i msty-*.rpm. With AppImage, libfuse2 may be required depending on the distribution.
  5. 05
    First launch
    Msty creates its data folder (~/Library/Application Support/Msty on Mac, %APPDATA%\Msty on Windows, ~/.config/Msty on Linux). Models, conversations, and Knowledge Stacks live there.

#2. First local model

Msty offers onboarding that downloads a small starter model with one click. You can also browse its catalog directly. Three solid choices in French to get started, depending on your hardware:

Small setup (8 GB RAM, no GPU)
Qwen 3.5 2B Q4_K_M (≈1.9 GB) or Granite 4.2 3B (≈2.2 GB). CPU-only is usable, ideal for chatting, brainstorming, and summarizing a short text.
Standard configuration (16 GB RAM, 6–8 GB GPU)
Qwen 3.5 9B Q4_K_M (≈6.6 GB), with 256k context and vision. The best versatile FR/EN everyday compromise.
Comfortable setup (32 GB RAM, 12–24 GB GPU)
Mistral Small 24B (≈14 GB) or Qwen 3.8 27B in Q4_K_M (≈18 GB). Significantly better reasoning, approaching the cloud experience.

To download: open the Local AI Models tab in the left sidebar, then click Browse and Download Online Model. Msty points to Hugging Face with a filter for compatible GGUF versions. Progress is displayed, and the model can be selected as soon as the download finishes.

→
Reuse your Ollama models
If you already have Ollama installed on the machine, Msty automatically detects the models present and makes them available through its interface—without duplicating the files. It's a low-key entry point for gradually migrating from Ollama CLI to a proper GUI.

#3. Cloud + local side by side

That's one of Msty's strengths: add your cloud API keys (OpenAI, Anthropic, Google, OpenRouter, Groq, Mistral cloud, etc.) and use these models in the same interface as your local models. Selecting the model for a new conversation is a simple dropdown that combines both worlds.

  1. 01
    Add a cloud provider
    Settings → Remote Model Providers → Add. Paste your API key. Msty stores the key locally (encrypted by the OS keychain when available).
  2. 02
    Choose the model for conversation
    In the chat bar, the selector displays all your models: Local (with a home icon), Cloud (with the provider's icon). You can switch models in the middle of a conversation; Msty preserves the context.
  3. 03
    Define rules
    Settings → Default Model lets you set the default model. Tip: setting a local model as the default forces you to choose explicitly when you want to send something to the cloud—useful for privacy.
!
Privacy: beware the cloud reflex
The convenience of the unified selector has a downside: it’s easy to accidentally send a sensitive prompt to OpenAI or Anthropic. If you handle confidential data, create a dedicated workspace with no cloud provider configured (see the next section), or remove the API keys for this context.

#4. Custom workspaces

A workspace in Msty is an isolated space with its own conversations, its own Knowledge Stacks, its own default model, and its own enabled providers. In practice, create one workspace per use case.

Perso
Local 7B model, no cloud, Knowledge Stacks over your personal Markdown notes, prompts in an informal tone.
Work
Local 14B model + cloud Claude for heavy tasks, Knowledge Stack on internal documentation (procedures, runbooks).
R&D / sandbox
Several models enabled for experimentation, frequent split chats, no Knowledge Stack.
Confidential
No cloud provider, local model only, OS-level encrypted Knowledge Stack, for processing customer or HR data.

Creation: left sidebar → Workspaces icon → New Workspace. Each workspace has its own icon and color, which helps prevent choosing the wrong context (and therefore the wrong model/destination).

#5. Knowledge Stacks (native RAG)

Knowledge Stacks are Msty's flagship feature. It's a complete RAG system with no configuration required: Msty handles chunking, generates embeddings with a local model (mxbai-embed-large by default), stores them in an embedded vector database, and automatically injects relevant passages into the LLM's context when you ask a question.

  1. 01
    Create a Knowledge Stack
    Sidebar → Knowledge Stacks → New Stack. Give it a name (e.g., “Legal docs”).
  2. 02
    Add sources
    Drag and drop files (PDF, DOCX, MD, TXT, CSV, JSON) or entire folders. You can also add URLs (Msty crawls the page), YouTube transcripts, or even a GitHub repository via URL.
  3. 03
    Wait for indexing
    Msty displays the progress: text extraction, chunking (≈500 tokens per chunk by default), and embedding generation. For 100 PDFs, expect 5-15 minutes depending on the GPU.
  4. 04
    Enable in a conversation
    In the chat, click the Knowledge Stack icon at the bottom → check the stack to use. You can activate several simultaneously. Msty displays the sources used below each response, with a direct link to the original passage.
→
Embedding model and French quality
By default, Msty uses mxbai-embed-large, which works reasonably well in French. For strictly French content, you can improve quality by switching to a French model such as Solon-embeddings or the multilingual bge-m3. Setting: Settings → Knowledge Stacks → Embedding Model. Everything remains local.
i
Limitations to know
Msty's RAG works very well for 10–500 documents. Beyond that—10k+ documents, BM25 hybrid search, custom reranking—a dedicated stack such as Qdrant + LlamaIndex or AnythingLLM remains more flexible. Msty optimizes the experience, not cardinality.

#6. Split Chat and comparisons

Split Chat lets you ask 2 models (free) or up to 4 models (Aurum) the same question in the same window, displayed in side-by-side columns. Ideal for calibrating a new model or choosing between two candidates without rewriting the prompt.

  1. 01
    Enable Split Chat
    In a conversation, click the Split button (columns icon) in the top bar. A second column appears with its own model selector.
  2. 02
    Configure each column
    Each column has its own model, temperature, and system prompt. You can compare a local Qwen 3.5 9B vs. Claude Haiku cloud on the same questions.
  3. 03
    Synchronized or independent conversation
    By default, every message you type is sent to both columns. You can uncheck Sync to let the conversations diverge (for example, if a response sparked an idea you want to explore in only one branch).
Useful comparison example
Prompt :
"Rédige un email professionnel poli pour refuser une réunion non urgente."

Colonne 1 : Qwen 3.5 9B local, température 0.7
Colonne 2 : Mistral Small 24B local, température 0.7

→ Comparaison directe du niveau de français, du ton, de la longueur,
de la pertinence des formules. Décide quel modèle
devient le défaut pour l'usage "rédaction".

#Msty vs Open WebUI vs LM Studio

All three are graphical interfaces for LLMs, but their center of gravity differs.

Msty
All-in-one desktop app, closed-source freemium. 1-click installation. Best for: no-code RAG (Knowledge Stacks), mixed cloud + local setups, workspaces, split chat. Less flexible for multi-user use or server deployment.
Open WebUI
Self-hosted web app, open source (BSD), Docker recommended. Native multi-user support (auth, roles), configurable RAG, accessible from any browser on the network. Best for: team/family deployment, full control, server hosting.
LM Studio
Closed-source but free desktop app. The richest Hugging Face catalog, native MLX for Mac, MCP support, and multi-token prediction. Best for: testing many models, raw performance, and advanced settings. Simple RAG (Chat with Documents), but less integrated than Knowledge Stacks.
→
How to choose between the three
You want local chat + RAG without coding and a polished UX: Msty. You want to deploy for multiple users on the network: Open WebUI. You want to explore models and get the most out of the inference engine: LM Studio. There's nothing stopping you from combining them—Msty can consume a Ollama served by Open WebUI on another machine.

#Troubleshooting

The model won't load (OOM)
Insufficient VRAM/RAM. Settings → Local AI → [model] → Advanced: lower n_gpu_layers (offload fewer layers to the GPU) or n_ctx (context window). Otherwise, switch to a smaller quantization (Q4_K_S, IQ3_M).
Low tokens/sec
Check GPU usage: nvidia-smi (NVIDIA), rocm-smi (AMD), or Activity Monitor GPU (Mac). If the GPU is at 0%, Msty is running on the CPU—check the drivers and make sure you are using the Local + GPU variant on Windows.
Knowledge Stack remains at 0%
Indexing requires the embeddings model to be downloaded. Settings → Knowledge Stacks → Embedding Model: make sure mxbai-embed-large (or your chosen model) is Downloaded; otherwise, click Download.
Cloud provider fails with 401
Invalid or expired API key. Check the provider's dashboard. For OpenAI, make sure you have active credits (keys without credit return 401 or 429).
Linux: AppImage won't launch
Install libfuse2 (sudo apt install libfuse2 on Ubuntu 22.04+). Launch the AppImage from a terminal to see the precise error.
Conversations lost after update
The data is in the Msty folder (Application Support on Mac, AppData on Windows, .config on Linux). Make sure no system cleaner has emptied this folder. Regularly backing up this folder is good practice.

#Go further

Depending on the direction you want to take:

Comparing Msty with its alternatives
“Ollama vs LM Studio vs Jan vs GPT4All” and “LM Studio vs Ollama: which should you choose in 2026?” provide the complete overview of local interfaces.
Taking RAG beyond Knowledge Stacks
“Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” covers the alternatives when you exceed the capabilities of Msty’s built-in RAG.
Choosing the right model for Msty
“Choosing your quantization (Q4, Q5, Q8, FP16)” helps weigh the quality/VRAM tradeoff in Msty's catalog.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.