Turn LM Studio into an OpenAI API server (2026)
In LM Studio, open the Developer tab and enable the Start server switch: the server listens on port 1234 and exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/responses, /v1/embeddings, /v1/models). Any OpenAI client works by simply changing the base URL. By default, it requires no authentication and listens only on localhost: API tokens and network access are configured in Server Settings.
LM Studio is more than a chat interface: its local server replaces OpenAI's API for your scripts, code editors, and agents, without sending a single line of text outside. This guide covers activation, every server setting, network access with authentication, on-demand model loading, and the real limits of a single workstation, taking the changes in version 0.4 into account.
#What you get
By the end of this guide, you’ll have an http://localhost:1234/v1 endpoint that any OpenAI SDK (Python, JavaScript, C#), LangChain, or coding tool such as Cline or Continue can use instead of the OpenAI API. The server also provides a native API at /api/v1 (stateful chat, model loading and downloading) and Anthropic-compatible endpoints. Everything stays on your machine—the model, requests, and responses do not leave your computer unless you expose it on the network yourself.
#1. Start the server
Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.
- Lifetime online access
- PDF + files
- Lifetime updates
- 01Open the Developer tabIn LM Studio, the Developer tab contains the server, its logs, and its settings. The model to be served must be downloaded beforehand.
- 02Enable Start serverToggle the Start server switch: the server starts on the port specified in Server Settings, 1234 in the documentation examples.
- 03Or start from the command lineFrom a terminal, the lms server start command launches the same server without opening the interface.
- 04Check the model listQuery /v1/models to confirm that the server responds and see the model identifiers to use in your requests.
#2. Server settings, one by one
The settings are under Developer, Server Settings. They determine who can call the server and what clients can make it do. Most integration problems come from one of these switches, usually the network or CORS setting.
| Setting | Role | Recommendation |
|---|---|---|
| Server Port | Server listening port (1234 in the documentation) | Change it if the port is already in use |
| Require Authentication | Requires a valid API token in the Authorization header | Enable it as soon as the server is no longer running on localhost |
| Serve on Local Network | Makes the server accessible to other devices on the local network | Disabled by default; combine with authentication |
| Allow per-request MCPs | Allows clients to use ephemeral remote MCP servers | Leave disabled unless you have a specific need |
| Allow calling servers from mcp.json | Allows clients to use the MCP servers defined in LM Studio | Requires authentication; risky if an MCP accesses your files |
| Enable CORS | Allows web applications from other origins | Only for a web application or certain extensions |
| Just in Time Model Loading | Loads models on demand, when the request arrives | Useful with third-party tools; see the dedicated section |
| Auto Unload Unused JIT Models | Unloads unused JIT models | Frees up memory |
| Only Keep Last JIT Loaded Model | Keep only the last model loaded on demand | Useful on a card with limited VRAM |
#3. Test with curl
A first control call is enough to validate the server. In requests, the model field must contain the model identifier as it appears in LM Studio, not a generic name such as local-model: the documentation reiterates this in its curl example.
The response is JSON in OpenAI format: choices[0].message.content contains the text. If the model field is wrong, or on-demand loading is disabled and the model isn't loaded, the request fails: first check the identifier returned by /v1/models.
#4. Call from Python
The openai SDK is used by changing only the base URL: this is the change shown in the LM Studio documentation. The SDK requires a key; as long as authentication is disabled, LM Studio does not check it, and if you enable authentication, the key becomes your API token.
Streaming works like OpenAI's: no client-side changes are needed to display tokens live. For agents and editors, LM Studio also implements the /v1/responses endpoint, described below.
#Which API to choose: OpenAI, Anthropic, or native
LM Studio exposes three endpoint families. OpenAI-compatible endpoints cover models, responses, chat, embeddings, and completions. Anthropic-compatible endpoints accept Anthropic’s message format. Since version 0.4.0, the native /api/v1 API adds functions specific to LM Studio: stateful chat, model loading, unloading, and downloading, and per-request context configuration.
| Need | Endpoint | Note |
|---|---|---|
| Replacing the OpenAI API in an existing tool | /v1/chat/completions | Streaming and custom tools supported |
| Codex-type agent or client | /v1/responses | Stateful chat and MCP available |
| Embeddings for a RAG | /v1/embeddings | Preloaded embedding model |
| Clients that speak the Anthropic format | Anthropic-compatible endpoints | Same server, different message format |
| Load, unload, and download a model | /api/v1/models/* | Native API, recommended by LM Studio since 0.4.0 |
| Set the context in the request | /api/v1/chat | The only endpoint that accepts context per request |
#6. Multiple models: on-demand loading and TTL
With just-in-time loading (JIT, for Just in Time), the first call to a model loads it into memory, and /v1/models lists all downloaded models, not just the ones that are loaded. Without JIT, /v1/models returns only models that are already loaded, and you must load the model before calling it. This mode is ideal when a tool such as Zed, Cline, or Continue chooses its own model.
- Default TTL
- A model loaded on demand is unloaded after 60 minutes without a request.
- TTL per request
- Add a ttl field (in seconds) to the request; 300 corresponds to 5 minutes.
- TTL for lms load
- Models loaded with lms load have no default TTL: use the --ttl option.
- Auto-Evict
- Enabled by default: only one model loaded on demand remains in memory at a time. Disable it to keep several models loaded.
#7. Expose the server on the network with authentication
For another computer on the network to call your server, enable Serve on Local Network in Server Settings, or start it with the listening address 0.0.0.0. The server will then stop listening on localhost alone: the documentation warns that any bind other than 127.0.0.1 exposes it beyond the machine and recommends enabling authentication.
Contrary to popular belief, LM Studio can authenticate requests. By default, it requires no authentication; when you enable the toggle in Server Settings, it accepts only requests carrying a valid API token, created in Manage Tokens with selected permissions. The token is displayed only when created: copy it immediately. This feature requires LM Studio 0.4.0 or later.
For access from the Internet, don’t expose the port: use a VPN or a reverse proxy with TLS. The principle is identical to that of a Ollama server, detailed in the security guide. A simpler alternative for using a model from another machine is LM Link, which serves a model from a remote device as if it were loaded locally.
#No graphical interface: llmster and automatic startup
Since version 0.4.0, the core of LM Studio has existed as a standalone daemon, llmster, designed to run without an interface on a Linux server, a GPU machine, or a local workstation. It installs with one command line, starts with lms daemon up, and the server then starts with lms server start. On a workstation with an interface, you can also check the option in the application settings that starts the server when you log in: closing the application then minimizes it to the system tray while the server continues running.
#9. Performance: what really matters
- GPU offload
- Load as many layers as possible into VRAM. A model that spills into system memory loses most of its throughput.
- Context Length
- Choose the context you need, not the maximum: the context cache uses VRAM and grows with the context length.
- Max Concurrent Predictions
- Number of requests processed simultaneously by a model; beyond that, they wait in a queue.
- Unified KV Cache
- Enabled by default: resources are not divided into fixed shares between requests, allowing variable request sizes.
The Flash Attention guide and the context-window guide detail the effects on memory. For substantial multi-user throughput, a dedicated server is still better suited: the vLLM guide shows the deployment.
#Limits and alternatives: what changed
Several commonly cited limitations are no longer true, and correcting them changes the choice of tool. The table compares what you still read with what the current documentation says.
| Common misconception | Reality |
|---|---|
| No authentication | API tokens available since 0.4.0, disabled by default |
| The requests run sequentially | Version 0.4.0 handles parallel requests to the same model (continuous batching), up to Max Concurrent Predictions; subsequent requests wait |
| Commercial license required for work | Free at home and at work since July 2025, according to LM Studio's announcement |
| Impossible without a graphical interface | llmster runs as a daemon, without a GUI |
| One workstation, no sharing | Serve on Local Network and LM Link let you serve other devices |
Real limitations remain: LM Studio is designed for a workstation, not a cluster; continuous batching does not replace a server designed for dozens of users, such as vLLM; and application updates can change behavior, so you must pin the version on a machine that provides a service. To choose between LM Studio and its competitors, compare them before committing.
- Deploy vLLM in production
- Ollama vs. LM Studio vs. Jan vs. GPT4All
- LM Studio on Linux
- Connect Cline to a local model
- Source: LM Studio documentation, local API server
- Source: LM Studio documentation, server settings
- Source: LM Studio documentation, authentication
- Source: LM Studio documentation, OpenAI compatibility
- Source: LM Studio 0.4.0 announcement
How do you enable the API server in LM Studio?+
How do you make the LM Studio server accessible from another PC?+
Does LM Studio provide API authentication?+
Does LM Studio process multiple requests in parallel?+
What identifier should you put in the model field?+
Is LM Studio free for business use?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.