Intermediate 12 minLM Studio

LM Studio Bionic: what is it and what is it for ?

LM Studio Bionic is the agent application that LM Studio split off from its standard app in July 2026. While the “historical” LM Studio remains focused on chat and the API server, Bionic adds agentic projects (Code and Work), document analysis (PDFs, spreadsheets, Word), and a Voxtral voice keyboard with 100% local transcription. This guide covers the app, explains what stays on your machine and what may be sent to the cloud, and details the settings that matter.

By Mohamed Meguedmi·Update 2026-09-05·Tested on Windows, macOS, and Linux
i
In brief
LM Studio Bionic is the agent application that LM Studio split off from the standard app in July 2026. · It adds Code (development) and Work (PDF, Excel, and Word document analysis) projects, plus a Voxtral voice keyboard for 100% local transcription. · A 7B to 14B model in Q4_K_M is enough to operate the agent (about 5 to 9 GB of VRAM). · Device mode keeps everything local; cloud mode offloads inference and sends data out—verify this before processing any sensitive content.

#What is LM Studio Bionic?

Bionic is a separate application developed by the same team as LM Studio. Rather than piling agentic features into the existing chat app, the publisher chose to create a separate product focused on “doing things” rather than “chatting.” The idea is a local agent that can read your files, write code in a project, and maintain persistent working context without sending your data to a third-party service by default.

In practical terms, the keyword lm studio bionic refers to this agent app: you open a project, give it a task (analyzing a folder of PDFs, refactoring a piece of code, summarizing a spreadsheet), and the local model executes steps by chaining file reading, reasoning, and writing. The inference engine remains the one you know from LM Studio (llama.cpp / MLX under the hood), with your already-downloaded GGUF or MLX models.

i
Separate app, shared engine
Bionic does not replace LM Studio. The two apps can coexist and share the same model folder. You can continue downloading your GGUF files into the standard LM Studio; Bionic reuses them for its agentic tasks.

#Bionic vs classic LM Studio: what’s the difference?

The Local Agents Kit

You just watched an agent run locally. The Local Agents Kit teaches you how to build one that really takes action (ch. 1), constrain it before unleashing it on your files (ch. 4), and measure whether it is reliable or merely creating an illusion (ch. 14).

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The most common mistake is thinking Bionic is simply an update to LM Studio. They are two apps with distinct goals. Here's how to tell them apart.

Classic LM Studio
Local chat, model management and downloads (GGUF/MLX from Hugging Face), OpenAI-compatible server, MCP support. This is your model hub and inference server.
Bionic
Agent app: persistent projects (Code and Work), document analysis, Voxtral voice keyboard, and multi-step task execution. Designed for producing work, not conversation.
Common ground
Same team, same local inference engine, same model library. You can serve a model from LM Studio and work in Bionic.
When to use which
LM Studio for chatting, testing a model, or exposing an API. Bionic when you want an agent to read your files and produce a concrete result.
→
The right reflex
If your need is “I want to ask a model questions,” stick with classic LM Studio. If it is “I want it to work on MY documents or MY code,” choose Bionic.

#Requirements and installation

Bionic runs on the same platforms as LM Studio (Windows, macOS Apple Silicon, Linux). Hardware requirements depend mainly on the model you will run locally to control the agent.

Recommended agent model
A 7B to 14B instruct model in Q4_K_M is enough for most tasks. Expect ≈5 GB of VRAM for a 7B model and ≈9 GB for a 14B model.
For serious code
A 32B (≈19 GB in Q4) or a Qwen3 Coder–style MoE delivers significantly better results for multi-file reasoning.
Reference hardware
RTX 3060 12GB or 4070 12GB for 7B–14B models; RTX 4090 24GB or Mac M4 Pro (24–48 GB unified) for 32B and larger.
Speech transcription
The Voxtral keyboard runs locally using a small speech-to-text model; allow for a few hundred MB to an additional 1–2 GB of memory.
  1. 01
    1. Download Bionic
    Go to the official LM Studio website (lmstudio.ai) and download the Bionic installer for your OS. This is a separate download from the classic LM Studio app.
  2. 02
    2. Point to your models
    On first launch, Bionic detects LM Studio's model directory. If you're starting from scratch, first download an instruct model into the standard LM Studio, then restart Bionic.
  3. 03
    3. Choose the agent model
    In Bionic’s settings, select the local model that will power your projects. Start with a 7B–14B Q4_K_M to verify that everything works before moving up the range.
  4. 04
    4. Check device mode
    By default, keep inference in device mode (local) to remain 100% offline. We explain device vs. cloud in more detail below.

#Code vs Work projects

Projects are at the heart of Bionic. A project groups a persistent work context: files, a task history, and agent behavior adapted to the type of work. Bionic distinguishes two families.

Code Project
Designed for development: the agent reads a codebase, understands the directory structure, writes and modifies files, and runs refactoring or code-generation tasks. It's the local equivalent of a coding agent, without sending your repo to the cloud.
Work project
Designed for office work and analysis: the agent processes documents (PDFs, Word files, spreadsheets), produces summaries, extracts data, and drafts reports. Ideal for monitoring, preparing case files, or document analysis.
Persistent context
Each project keeps its history and attached files. You can close Bionic and resume where you left off, without providing the context again in every session.

The Code/Work distinction is not merely aesthetic: it changes which tools the agent is allowed to use and how it structures its responses. A Code project favors file editing and execution; a Work project favors reading, extraction, and writing.

→
One project = one goal
Avoid the catch-all project. One project per codebase or document folder keeps the context clean and significantly improves agent relevance, especially with a modest-sized local model.

#Analyzing documents, PDFs, and spreadsheets

That’s the strength of Work projects. You drag files into the project, and the agent reads them to answer your requests. Bionic natively supports several common office formats.

PDF
Text extraction (and structure extraction when it’s clean) for summarizing, querying, or comparing multiple documents.
Spreadsheets
Reading Excel/CSV sheets: the agent can describe columns, calculate simple aggregates, detect anomalies, or produce a numerical summary.
Word / text documents
Summaries, rewrites, key information extraction, and note formatting.

Processing happens through the local model: your documents don't leave the machine as long as you stay in device mode. That's the major difference from a cloud assistant—a confidential folder (contracts, HR data, health data) can be analyzed without exfiltration.

!
Watch out for the context window
A large PDF or spreadsheet can exceed the context window of a 7B model. If responses become incoherent, split the document, increase the model's context size, or switch to a model with a larger window. Analysis quality depends directly on the local model you choose.

#Voxtral voice keyboard: local speech transcription

Bionic includes a voice keyboard based on Voxtral, a speech-recognition model that runs locally. You dictate your request, it is transcribed into text on your machine, and then sent to the agent. No audio is sent to an online transcription service.

100% local
Transcription runs through the Voxtral model loaded in memory. Your voice is sent to no server, making it suitable for sensitive content or disconnected environments.
Use cases
Dictate a long instruction to the agent, take voice notes in a Work project, or describe a coding task without letting go of the physical keyboard.
Languages
Voxtral supports French and English, among others; quality depends on the installed transcription model and your microphone.
i
Voxtral ≠ the chat model
Voxtral only converts speech to text. Your agent model (the 7B/14B/32B model you chose) then processes the transcribed text. Both models coexist in memory — keep that in mind when budgeting VRAM.

#Device vs. cloud mode: what stays local

Bionic offers two inference modes, and this is the most important setting to understand if you want to remain in control of your data.

Device mode (local)
Inference runs on your machine, using your GGUF/MLX models. Nothing leaves it: not your files, not your prompts, not your voice. This is the mode to choose for strict self-hosting.
Cloud mode
Inference is offloaded to a remote machine, which is useful for running models that are too large for your hardware. Convenient, but your data then leaves the machine—reserve this for non-sensitive content.
Choice by project
You can keep a confidential Work project on-device and a less sensitive Code project in the cloud if you need a more powerful model. Check the active mode before loading sensitive documents.
!
Check the mode before handling sensitive material
Cloud mode changes the privacy equation. Before dropping a contract or personal data into a project, confirm that you're in device mode. That's the only guarantee that nothing is sent online.

#Settings that matter and local footprint

A few parameters make the difference between a slow, frustrating agent and a responsive, reliable one.

Agent model
The number-one choice. Too small, and the agent gets lost in multi-step tasks; too large, and it crawls. A 14B Q4_K_M is a good balance for getting started.
Context size
Increase it if you analyze long documents, but keep an eye on VRAM—a larger context uses more memory.
Quantization
Q4_K_M for a balance of speed and quality, Q5_K_M or Q8_0 if you have the VRAM and want greater accuracy on code.
GPU offloading
Push as many layers as possible onto the GPU for speed. If the model spills over, reduce the quantization or model size instead of offloading everything to the CPU.

What remains local: in device mode, the entire pipeline — Voxtral transcription, reading your files, and agent inference — runs on your machine. Bionic stays true to the LM Studio ethos: your models, your data, your hardware.


#Troubleshooting

Bionic can't see my models
Make sure the model directory points to LM Studio. Download at least one instruct model into the standard LM Studio, then restart Bionic.
The agent is very slow
Model too large for your VRAM: it offloads to the CPU. Drop down one size (32B → 14B) or one quantization level (Q8 → Q4_K_M).
Inconsistent responses on a large document
The document exceeds the context window. Split it, increase the context size, or choose a model with a larger window.
Voice dictation won’t start
Check your OS microphone permissions and make sure the Voxtral model is installed correctly. Local transcription needs additional free memory beyond the agent model.
Privacy concerns
Open Settings and confirm the device mode. In the cloud, prompts and files are sent remotely.

#Go further

Bionic builds on the LM Studio ecosystem. These guides supplement the onboarding and help you choose the right model to drive your agents.

Getting started with LM Studio
“LM Studio in 2026: complete tutorial” — install the classic app, download models, and configure the GPU, an essential foundation before Bionic.
Choose your quantization
“Choosing your quantization (Q4, Q5, Q8, FP16)” — to tune the speed/quality tradeoff of your agent model.
Local RAG
“Local RAG with LM Studio: chat with your documents” — for a deeper look at document analysis beyond Work projects.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.