LM Studio Bionic: what is it and what is it for ?
LM Studio Bionic is the agent application that LM Studio split off from its standard app in July 2026. While the “historical” LM Studio remains focused on chat and the API server, Bionic adds agentic projects (Code and Work), document analysis (PDFs, spreadsheets, Word), and a Voxtral voice keyboard with 100% local transcription. This guide covers the app, explains what stays on your machine and what may be sent to the cloud, and details the settings that matter.
#What is LM Studio Bionic?
Bionic is a separate application developed by the same team as LM Studio. Rather than piling agentic features into the existing chat app, the publisher chose to create a separate product focused on “doing things” rather than “chatting.” The idea is a local agent that can read your files, write code in a project, and maintain persistent working context without sending your data to a third-party service by default.
In practical terms, the keyword lm studio bionic refers to this agent app: you open a project, give it a task (analyzing a folder of PDFs, refactoring a piece of code, summarizing a spreadsheet), and the local model executes steps by chaining file reading, reasoning, and writing. The inference engine remains the one you know from LM Studio (llama.cpp / MLX under the hood), with your already-downloaded GGUF or MLX models.
#Bionic vs classic LM Studio: what’s the difference?
You just watched an agent run locally. The Local Agents Kit teaches you how to build one that really takes action (ch. 1), constrain it before unleashing it on your files (ch. 4), and measure whether it is reliable or merely creating an illusion (ch. 14).
- Lifetime online access
- PDF + files
- Lifetime updates
The most common mistake is thinking Bionic is simply an update to LM Studio. They are two apps with distinct goals. Here's how to tell them apart.
- Classic LM Studio
- Local chat, model management and downloads (GGUF/MLX from Hugging Face), OpenAI-compatible server, MCP support. This is your model hub and inference server.
- Bionic
- Agent app: persistent projects (Code and Work), document analysis, Voxtral voice keyboard, and multi-step task execution. Designed for producing work, not conversation.
- Common ground
- Same team, same local inference engine, same model library. You can serve a model from LM Studio and work in Bionic.
- When to use which
- LM Studio for chatting, testing a model, or exposing an API. Bionic when you want an agent to read your files and produce a concrete result.
#Requirements and installation
Bionic runs on the same platforms as LM Studio (Windows, macOS Apple Silicon, Linux). Hardware requirements depend mainly on the model you will run locally to control the agent.
- Recommended agent model
- A 7B to 14B instruct model in Q4_K_M is enough for most tasks. Expect ≈5 GB of VRAM for a 7B model and ≈9 GB for a 14B model.
- For serious code
- A 32B (≈19 GB in Q4) or a Qwen3 Coder–style MoE delivers significantly better results for multi-file reasoning.
- Reference hardware
- RTX 3060 12GB or 4070 12GB for 7B–14B models; RTX 4090 24GB or Mac M4 Pro (24–48 GB unified) for 32B and larger.
- Speech transcription
- The Voxtral keyboard runs locally using a small speech-to-text model; allow for a few hundred MB to an additional 1–2 GB of memory.
- 011. Download BionicGo to the official LM Studio website (lmstudio.ai) and download the Bionic installer for your OS. This is a separate download from the classic LM Studio app.
- 022. Point to your modelsOn first launch, Bionic detects LM Studio's model directory. If you're starting from scratch, first download an instruct model into the standard LM Studio, then restart Bionic.
- 033. Choose the agent modelIn Bionic’s settings, select the local model that will power your projects. Start with a 7B–14B Q4_K_M to verify that everything works before moving up the range.
- 044. Check device modeBy default, keep inference in device mode (local) to remain 100% offline. We explain device vs. cloud in more detail below.
#Code vs Work projects
Projects are at the heart of Bionic. A project groups a persistent work context: files, a task history, and agent behavior adapted to the type of work. Bionic distinguishes two families.
- Code Project
- Designed for development: the agent reads a codebase, understands the directory structure, writes and modifies files, and runs refactoring or code-generation tasks. It's the local equivalent of a coding agent, without sending your repo to the cloud.
- Work project
- Designed for office work and analysis: the agent processes documents (PDFs, Word files, spreadsheets), produces summaries, extracts data, and drafts reports. Ideal for monitoring, preparing case files, or document analysis.
- Persistent context
- Each project keeps its history and attached files. You can close Bionic and resume where you left off, without providing the context again in every session.
The Code/Work distinction is not merely aesthetic: it changes which tools the agent is allowed to use and how it structures its responses. A Code project favors file editing and execution; a Work project favors reading, extraction, and writing.
#Analyzing documents, PDFs, and spreadsheets
That’s the strength of Work projects. You drag files into the project, and the agent reads them to answer your requests. Bionic natively supports several common office formats.
- Text extraction (and structure extraction when it’s clean) for summarizing, querying, or comparing multiple documents.
- Spreadsheets
- Reading Excel/CSV sheets: the agent can describe columns, calculate simple aggregates, detect anomalies, or produce a numerical summary.
- Word / text documents
- Summaries, rewrites, key information extraction, and note formatting.
Processing happens through the local model: your documents don't leave the machine as long as you stay in device mode. That's the major difference from a cloud assistant—a confidential folder (contracts, HR data, health data) can be analyzed without exfiltration.
#Voxtral voice keyboard: local speech transcription
Bionic includes a voice keyboard based on Voxtral, a speech-recognition model that runs locally. You dictate your request, it is transcribed into text on your machine, and then sent to the agent. No audio is sent to an online transcription service.
- 100% local
- Transcription runs through the Voxtral model loaded in memory. Your voice is sent to no server, making it suitable for sensitive content or disconnected environments.
- Use cases
- Dictate a long instruction to the agent, take voice notes in a Work project, or describe a coding task without letting go of the physical keyboard.
- Languages
- Voxtral supports French and English, among others; quality depends on the installed transcription model and your microphone.
#Device vs. cloud mode: what stays local
Bionic offers two inference modes, and this is the most important setting to understand if you want to remain in control of your data.
- Device mode (local)
- Inference runs on your machine, using your GGUF/MLX models. Nothing leaves it: not your files, not your prompts, not your voice. This is the mode to choose for strict self-hosting.
- Cloud mode
- Inference is offloaded to a remote machine, which is useful for running models that are too large for your hardware. Convenient, but your data then leaves the machine—reserve this for non-sensitive content.
- Choice by project
- You can keep a confidential Work project on-device and a less sensitive Code project in the cloud if you need a more powerful model. Check the active mode before loading sensitive documents.
#Settings that matter and local footprint
A few parameters make the difference between a slow, frustrating agent and a responsive, reliable one.
- Agent model
- The number-one choice. Too small, and the agent gets lost in multi-step tasks; too large, and it crawls. A 14B Q4_K_M is a good balance for getting started.
- Context size
- Increase it if you analyze long documents, but keep an eye on VRAM—a larger context uses more memory.
- Quantization
- Q4_K_M for a balance of speed and quality, Q5_K_M or Q8_0 if you have the VRAM and want greater accuracy on code.
- GPU offloading
- Push as many layers as possible onto the GPU for speed. If the model spills over, reduce the quantization or model size instead of offloading everything to the CPU.
What remains local: in device mode, the entire pipeline — Voxtral transcription, reading your files, and agent inference — runs on your machine. Bionic stays true to the LM Studio ethos: your models, your data, your hardware.
#Troubleshooting
- Bionic can't see my models
- Make sure the model directory points to LM Studio. Download at least one instruct model into the standard LM Studio, then restart Bionic.
- The agent is very slow
- Model too large for your VRAM: it offloads to the CPU. Drop down one size (32B → 14B) or one quantization level (Q8 → Q4_K_M).
- Inconsistent responses on a large document
- The document exceeds the context window. Split it, increase the context size, or choose a model with a larger window.
- Voice dictation won’t start
- Check your OS microphone permissions and make sure the Voxtral model is installed correctly. Local transcription needs additional free memory beyond the agent model.
- Privacy concerns
- Open Settings and confirm the device mode. In the cloud, prompts and files are sent remotely.
#Go further
Bionic builds on the LM Studio ecosystem. These guides supplement the onboarding and help you choose the right model to drive your agents.
- Getting started with LM Studio
- “LM Studio in 2026: complete tutorial” — install the classic app, download models, and configure the GPU, an essential foundation before Bionic.
- Choose your quantization
- “Choosing your quantization (Q4, Q5, Q8, FP16)” — to tune the speed/quality tradeoff of your agent model.
- Local RAG
- “Local RAG with LM Studio: chat with your documents” — for a deeper look at document analysis beyond Work projects.
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.