Msty: the elegant local LLM interface (Mac, Windows, Linux)
Msty is a desktop app focused entirely on the user experience for running LLMs locally. Where Open WebUI requires Docker and LM Studio focuses on the engine, Msty offers everything in a single installer: a local model, cloud models (OpenAI, Claude, Gemini) side by side, native RAG through Knowledge Stacks, workspaces, and split chat for comparing two models. This guide shows how to install and use Msty on Mac, Windows, and Linux for serious local LLM use.
#Why Msty
Msty (msty.app) targets a specific profile: someone who wants a clean graphical app, without Docker or a CLI, but with real features beyond simple chat. Three differentiators sum up the tool.
- Everything in a single app
- Built-in inference engine (based on llama.cpp), client for Ollama if you already installed it, and cloud connectors—no stack to assemble.
- Knowledge Stacks
- RAG is built in and requires no code. Drag and drop PDFs, Markdown, entire folders, or even URLs; Msty handles chunking, embeddings, and retrieval.
- Workspaces and Split Chat
- Several isolated workspaces (personal, work, R&D), with the ability to ask the same question to 2-4 models in parallel in the same view.
#Prerequisites
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- OS
- macOS 11+ (Apple Silicon or Intel), Windows 10/11 (x64), Linux x64 (AppImage and .deb/.rpm packages).
- Minimum RAM
- 8 GB for a 3B Q4_K_M model, 16 GB for 7B, 32 GB for a comfortable 14B, 64 GB for targeting 32B.
- GPU
- Optional but recommended. NVIDIA with CUDA, AMD with ROCm/Vulkan, Apple Silicon with Metal — Msty detects and uses them automatically.
- Storage
- Allow 10-50 GB depending on the number of models. GGUF Q4 files range from 2 GB (3B) to 40 GB (70B).
#1. Installation
Msty can only be downloaded from the official website. There is no package manager support yet (neither Homebrew nor winget), which makes it easier to verify the source.
- 01Download the installerGo to msty.app and choose your OS. Three variants are available: Online (reduced, cloud-only), Local (recommended, includes the inference engine), and Local + GPU (with CUDA drivers bundled for Windows).
- 02Install (Mac)Open the .dmg, drag Msty to Applications, then launch it for the first time with right-click → Open (Gatekeeper). macOS asks for confirmation because the app is not notarized by the App Store.
- 03Install (Windows)Run the .exe. Choose the GPU variant if you have a recent NVIDIA—it includes the CUDA runtimes and avoids manually wrestling with the drivers. Otherwise, the Local CPU/Vulkan version is sufficient.
- 04Install (Linux)AppImage: chmod +x Msty-*.AppImage, then ./Msty-*.AppImage. .deb: sudo dpkg -i msty_*.deb. .rpm: sudo rpm -i msty-*.rpm. With AppImage, libfuse2 may be required depending on the distribution.
- 05First launchMsty creates its data folder (~/Library/Application Support/Msty on Mac, %APPDATA%\Msty on Windows, ~/.config/Msty on Linux). Models, conversations, and Knowledge Stacks live there.
#2. First local model
Msty offers onboarding that downloads a small starter model with one click. You can also browse its catalog directly. Three solid choices in French to get started, depending on your hardware:
- Small setup (8 GB RAM, no GPU)
- Qwen 3.5 2B Q4_K_M (≈1.9 GB) or Granite 4.2 3B (≈2.2 GB). CPU-only is usable, ideal for chatting, brainstorming, and summarizing a short text.
- Standard configuration (16 GB RAM, 6–8 GB GPU)
- Qwen 3.5 9B Q4_K_M (≈6.6 GB), with 256k context and vision. The best versatile FR/EN everyday compromise.
- Comfortable setup (32 GB RAM, 12–24 GB GPU)
- Mistral Small 24B (≈14 GB) or Qwen 3.8 27B in Q4_K_M (≈18 GB). Significantly better reasoning, approaching the cloud experience.
To download: open the Local AI Models tab in the left sidebar, then click Browse and Download Online Model. Msty points to Hugging Face with a filter for compatible GGUF versions. Progress is displayed, and the model can be selected as soon as the download finishes.
#3. Cloud + local side by side
That's one of Msty's strengths: add your cloud API keys (OpenAI, Anthropic, Google, OpenRouter, Groq, Mistral cloud, etc.) and use these models in the same interface as your local models. Selecting the model for a new conversation is a simple dropdown that combines both worlds.
- 01Add a cloud providerSettings → Remote Model Providers → Add. Paste your API key. Msty stores the key locally (encrypted by the OS keychain when available).
- 02Choose the model for conversationIn the chat bar, the selector displays all your models: Local (with a home icon), Cloud (with the provider's icon). You can switch models in the middle of a conversation; Msty preserves the context.
- 03Define rulesSettings → Default Model lets you set the default model. Tip: setting a local model as the default forces you to choose explicitly when you want to send something to the cloud—useful for privacy.
#4. Custom workspaces
A workspace in Msty is an isolated space with its own conversations, its own Knowledge Stacks, its own default model, and its own enabled providers. In practice, create one workspace per use case.
- Perso
- Local 7B model, no cloud, Knowledge Stacks over your personal Markdown notes, prompts in an informal tone.
- Work
- Local 14B model + cloud Claude for heavy tasks, Knowledge Stack on internal documentation (procedures, runbooks).
- R&D / sandbox
- Several models enabled for experimentation, frequent split chats, no Knowledge Stack.
- Confidential
- No cloud provider, local model only, OS-level encrypted Knowledge Stack, for processing customer or HR data.
Creation: left sidebar → Workspaces icon → New Workspace. Each workspace has its own icon and color, which helps prevent choosing the wrong context (and therefore the wrong model/destination).
#5. Knowledge Stacks (native RAG)
Knowledge Stacks are Msty's flagship feature. It's a complete RAG system with no configuration required: Msty handles chunking, generates embeddings with a local model (mxbai-embed-large by default), stores them in an embedded vector database, and automatically injects relevant passages into the LLM's context when you ask a question.
- 01Create a Knowledge StackSidebar → Knowledge Stacks → New Stack. Give it a name (e.g., “Legal docs”).
- 02Add sourcesDrag and drop files (PDF, DOCX, MD, TXT, CSV, JSON) or entire folders. You can also add URLs (Msty crawls the page), YouTube transcripts, or even a GitHub repository via URL.
- 03Wait for indexingMsty displays the progress: text extraction, chunking (≈500 tokens per chunk by default), and embedding generation. For 100 PDFs, expect 5-15 minutes depending on the GPU.
- 04Enable in a conversationIn the chat, click the Knowledge Stack icon at the bottom → check the stack to use. You can activate several simultaneously. Msty displays the sources used below each response, with a direct link to the original passage.
#6. Split Chat and comparisons
Split Chat lets you ask 2 models (free) or up to 4 models (Aurum) the same question in the same window, displayed in side-by-side columns. Ideal for calibrating a new model or choosing between two candidates without rewriting the prompt.
- 01Enable Split ChatIn a conversation, click the Split button (columns icon) in the top bar. A second column appears with its own model selector.
- 02Configure each columnEach column has its own model, temperature, and system prompt. You can compare a local Qwen 3.5 9B vs. Claude Haiku cloud on the same questions.
- 03Synchronized or independent conversationBy default, every message you type is sent to both columns. You can uncheck Sync to let the conversations diverge (for example, if a response sparked an idea you want to explore in only one branch).
#Msty vs Open WebUI vs LM Studio
All three are graphical interfaces for LLMs, but their center of gravity differs.
- Msty
- All-in-one desktop app, closed-source freemium. 1-click installation. Best for: no-code RAG (Knowledge Stacks), mixed cloud + local setups, workspaces, split chat. Less flexible for multi-user use or server deployment.
- Open WebUI
- Self-hosted web app, open source (BSD), Docker recommended. Native multi-user support (auth, roles), configurable RAG, accessible from any browser on the network. Best for: team/family deployment, full control, server hosting.
- LM Studio
- Closed-source but free desktop app. The richest Hugging Face catalog, native MLX for Mac, MCP support, and multi-token prediction. Best for: testing many models, raw performance, and advanced settings. Simple RAG (Chat with Documents), but less integrated than Knowledge Stacks.
#Troubleshooting
- The model won't load (OOM)
- Insufficient VRAM/RAM. Settings → Local AI → [model] → Advanced: lower n_gpu_layers (offload fewer layers to the GPU) or n_ctx (context window). Otherwise, switch to a smaller quantization (Q4_K_S, IQ3_M).
- Low tokens/sec
- Check GPU usage: nvidia-smi (NVIDIA), rocm-smi (AMD), or Activity Monitor GPU (Mac). If the GPU is at 0%, Msty is running on the CPU—check the drivers and make sure you are using the Local + GPU variant on Windows.
- Knowledge Stack remains at 0%
- Indexing requires the embeddings model to be downloaded. Settings → Knowledge Stacks → Embedding Model: make sure mxbai-embed-large (or your chosen model) is Downloaded; otherwise, click Download.
- Cloud provider fails with 401
- Invalid or expired API key. Check the provider's dashboard. For OpenAI, make sure you have active credits (keys without credit return 401 or 429).
- Linux: AppImage won't launch
- Install libfuse2 (sudo apt install libfuse2 on Ubuntu 22.04+). Launch the AppImage from a terminal to see the precise error.
- Conversations lost after update
- The data is in the Msty folder (Application Support on Mac, AppData on Windows, .config on Linux). Make sure no system cleaner has emptied this folder. Regularly backing up this folder is good practice.
#Go further
Depending on the direction you want to take:
- Comparing Msty with its alternatives
- “Ollama vs LM Studio vs Jan vs GPT4All” and “LM Studio vs Ollama: which should you choose in 2026?” provide the complete overview of local interfaces.
- Taking RAG beyond Knowledge Stacks
- “Local RAG with Ollama without coding (Open WebUI, AnythingLLM)” covers the alternatives when you exceed the capabilities of Msty’s built-in RAG.
- Choosing the right model for Msty
- “Choosing your quantization (Q4, Q5, Q8, FP16)” helps weigh the quality/VRAM tradeoff in Msty's catalog.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.