Intermediate 11 minLM Studio

Turn LM Studio into an OpenAI API server (2026)

Direct response

In LM Studio, open the Developer tab and enable the Start server switch: the server listens on port 1234 and exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/responses, /v1/embeddings, /v1/models). Any OpenAI client works by simply changing the base URL. By default, it requires no authentication and listens only on localhost: API tokens and network access are configured in Server Settings.

LM Studio is more than a chat interface: its local server replaces OpenAI's API for your scripts, code editors, and agents, without sending a single line of text outside. This guide covers activation, every server setting, network access with authentication, on-demand model loading, and the real limits of a single workstation, taking the changes in version 0.4 into account.

By Mohamed Meguedmi·Update 2026-09-30·Tested on Windows, macOS, and Linux

#What you get

By the end of this guide, you’ll have an http://localhost:1234/v1 endpoint that any OpenAI SDK (Python, JavaScript, C#), LangChain, or coding tool such as Cline or Continue can use instead of the OpenAI API. The server also provides a native API at /api/v1 (stateful chat, model loading and downloading) and Anthropic-compatible endpoints. Everything stays on your machine—the model, requests, and responses do not leave your computer unless you expose it on the network yourself.

#1. Start the server

The Local AI Kit

Your private ChatGPT, free, on your own machine in an hour — LM Studio, Ollama, Open WebUI, your documents, no cloud.

  • Lifetime online access
  • PDF + files
  • Lifetime updates
  1. 01
    Open the Developer tab
    In LM Studio, the Developer tab contains the server, its logs, and its settings. The model to be served must be downloaded beforehand.
  2. 02
    Enable Start server
    Toggle the Start server switch: the server starts on the port specified in Server Settings, 1234 in the documentation examples.
  3. 03
    Or start from the command line
    From a terminal, the lms server start command launches the same server without opening the interface.
  4. 04
    Check the model list
    Query /v1/models to confirm that the server responds and see the model identifiers to use in your requests.
Start the server from the terminal
lms server start

#2. Server settings, one by one

The settings are under Developer, Server Settings. They determine who can call the server and what clients can make it do. Most integration problems come from one of these switches, usually the network or CORS setting.

LM Studio Server Settings (official documentation)
SettingRoleRecommendation
Server PortServer listening port (1234 in the documentation)Change it if the port is already in use
Require AuthenticationRequires a valid API token in the Authorization headerEnable it as soon as the server is no longer running on localhost
Serve on Local NetworkMakes the server accessible to other devices on the local networkDisabled by default; combine with authentication
Allow per-request MCPsAllows clients to use ephemeral remote MCP serversLeave disabled unless you have a specific need
Allow calling servers from mcp.jsonAllows clients to use the MCP servers defined in LM StudioRequires authentication; risky if an MCP accesses your files
Enable CORSAllows web applications from other originsOnly for a web application or certain extensions
Just in Time Model LoadingLoads models on demand, when the request arrivesUseful with third-party tools; see the dedicated section
Auto Unload Unused JIT ModelsUnloads unused JIT modelsFrees up memory
Only Keep Last JIT Loaded ModelKeep only the last model loaded on demandUseful on a card with limited VRAM
!
A setting that exposes your files
The option that allows calls to the servers in mcp.json gives API clients access to the tools you defined there. The documentation advises against it without authentication and in fact requires Require Authentication to be enabled. Enable it only if you understand the scope of every declared MCP server.

#3. Test with curl

A first control call is enough to validate the server. In requests, the model field must contain the model identifier as it appears in LM Studio, not a generic name such as local-model: the documentation reiterates this in its curl example.

Model list
curl http://localhost:1234/v1/models
Chat completion
curl http://localhost:1234/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "IDENTIFIANT-DU-MODELE",
    "messages": [
      {"role":"system","content":"Tu es concis."},
      {"role":"user","content":"Capitale du Portugal ?"}
    ],
    "temperature": 0.2
  }'

The response is JSON in OpenAI format: choices[0].message.content contains the text. If the model field is wrong, or on-demand loading is disabled and the model isn't loaded, the request fails: first check the identifier returned by /v1/models.

#4. Call from Python

The openai SDK is used by changing only the base URL: this is the change shown in the LM Studio documentation. The SDK requires a key; as long as authentication is disabled, LM Studio does not check it, and if you enable authentication, the key becomes your API token.

Via the openai SDK
from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:1234/v1",
    api_key="lm-studio",  # ignorée sans authentification ; votre jeton sinon
)

resp = client.chat.completions.create(
    model="IDENTIFIANT-DU-MODELE",
    messages=[
        {"role": "system", "content": "Réponds en 1 phrase."},
        {"role": "user",   "content": "Qu'est-ce qu'un LLM ?"},
    ],
    temperature=0.3,
    stream=True,
)

for chunk in resp:
    delta = chunk.choices[0].delta.content or ""
    print(delta, end="", flush=True)

Streaming works like OpenAI's: no client-side changes are needed to display tokens live. For agents and editors, LM Studio also implements the /v1/responses endpoint, described below.

#Which API to choose: OpenAI, Anthropic, or native

LM Studio exposes three endpoint families. OpenAI-compatible endpoints cover models, responses, chat, embeddings, and completions. Anthropic-compatible endpoints accept Anthropic’s message format. Since version 0.4.0, the native /api/v1 API adds functions specific to LM Studio: stateful chat, model loading, unloading, and downloading, and per-request context configuration.

Endpoints by need
NeedEndpointNote
Replacing the OpenAI API in an existing tool/v1/chat/completionsStreaming and custom tools supported
Codex-type agent or client/v1/responsesStateful chat and MCP available
Embeddings for a RAG/v1/embeddingsPreloaded embedding model
Clients that speak the Anthropic formatAnthropic-compatible endpointsSame server, different message format
Load, unload, and download a model/api/v1/models/*Native API, recommended by LM Studio since 0.4.0
Set the context in the request/api/v1/chatThe only endpoint that accepts context per request

#6. Multiple models: on-demand loading and TTL

With just-in-time loading (JIT, for Just in Time), the first call to a model loads it into memory, and /v1/models lists all downloaded models, not just the ones that are loaded. Without JIT, /v1/models returns only models that are already loaded, and you must load the model before calling it. This mode is ideal when a tool such as Zed, Cline, or Continue chooses its own model.

Default TTL
A model loaded on demand is unloaded after 60 minutes without a request.
TTL per request
Add a ttl field (in seconds) to the request; 300 corresponds to 5 minutes.
TTL for lms load
Models loaded with lms load have no default TTL: use the --ttl option.
Auto-Evict
Enabled by default: only one model loaded on demand remains in memory at a time. Disable it to keep several models loaded.
!
Two models, twice the weights
Auto-Evict unloads the previous model before loading a new one. If you disable it, the weights add up: an 8-billion-parameter model in Q4 (about 5 GB) and a 9-billion-parameter model (about 6 GB) together occupy more than 10 GB even before the context. Monitor your VRAM.

#7. Expose the server on the network with authentication

For another computer on the network to call your server, enable Serve on Local Network in Server Settings, or start it with the listening address 0.0.0.0. The server will then stop listening on localhost alone: the documentation warns that any bind other than 127.0.0.1 exposes it beyond the machine and recommends enabling authentication.

Listen on all IPv4 interfaces
lms server start --bind 0.0.0.0
Remote client
curl http://192.168.1.42:1234/v1/models

Contrary to popular belief, LM Studio can authenticate requests. By default, it requires no authentication; when you enable the toggle in Server Settings, it accepts only requests carrying a valid API token, created in Manage Tokens with selected permissions. The token is displayed only when created: copy it immediately. This feature requires LM Studio 0.4.0 or later.

Call with an API token
curl http://192.168.1.42:1234/v1/models \
  -H "Authorization: Bearer $LM_API_TOKEN"

For access from the Internet, don’t expose the port: use a VPN or a reverse proxy with TLS. The principle is identical to that of a Ollama server, detailed in the security guide. A simpler alternative for using a model from another machine is LM Link, which serves a model from a remote device as if it were loaded locally.

#No graphical interface: llmster and automatic startup

Since version 0.4.0, the core of LM Studio has existed as a standalone daemon, llmster, designed to run without an interface on a Linux server, a GPU machine, or a local workstation. It installs with one command line, starts with lms daemon up, and the server then starts with lms server start. On a workstation with an interface, you can also check the option in the application settings that starts the server when you log in: closing the application then minimizes it to the system tray while the server continues running.

Install and start llmster (Linux and Mac)
curl -fsSL https://lmstudio.ai/install.sh | bash
lms daemon up
lms server start

#9. Performance: what really matters

GPU offload
Load as many layers as possible into VRAM. A model that spills into system memory loses most of its throughput.
Context Length
Choose the context you need, not the maximum: the context cache uses VRAM and grows with the context length.
Max Concurrent Predictions
Number of requests processed simultaneously by a model; beyond that, they wait in a queue.
Unified KV Cache
Enabled by default: resources are not divided into fixed shares between requests, allowing variable request sizes.

The Flash Attention guide and the context-window guide detail the effects on memory. For substantial multi-user throughput, a dedicated server is still better suited: the vLLM guide shows the deployment.

#Limits and alternatives: what changed

Several commonly cited limitations are no longer true, and correcting them changes the choice of tool. The table compares what you still read with what the current documentation says.

Myths and reality (LM Studio documentation, 2026)
Common misconceptionReality
No authenticationAPI tokens available since 0.4.0, disabled by default
The requests run sequentiallyVersion 0.4.0 handles parallel requests to the same model (continuous batching), up to Max Concurrent Predictions; subsequent requests wait
Commercial license required for workFree at home and at work since July 2025, according to LM Studio's announcement
Impossible without a graphical interfacellmster runs as a daemon, without a GUI
One workstation, no sharingServe on Local Network and LM Link let you serve other devices

Real limitations remain: LM Studio is designed for a workstation, not a cluster; continuous batching does not replace a server designed for dozens of users, such as vLLM; and application updates can change behavior, so you must pin the version on a machine that provides a service. To choose between LM Studio and its competitors, compare them before committing.

FAQ
How do you enable the API server in LM Studio?+
Open the Developer tab and turn on the Start server switch, or run the command lms server start in a terminal. The server listens on the port configured in Server Settings, 1234 in the documentation. Test it with curl http://localhost:1234/v1/models, which returns the available models; if the list is empty, first load a model or enable on-demand loading in Server Settings.
How do you make the LM Studio server accessible from another PC?+
Enable Serve on Local Network in Server Settings, or start with lms server start --bind 0.0.0.0. The server will then listen beyond localhost: also enable API token authentication, because the documentation recommends this precaution for any bind other than 127.0.0.1. From another machine, use the computer's IP address and the same port, for example http://192.168.1.42:1234/v1.
Does LM Studio provide API authentication?+
Yes, since version 0.4.0: API tokens are created in Manage Tokens and enabled with Require Authentication in Server Settings. By default, no authentication is required. Once enabled, all REST requests and SDK requests must include a valid token in the Authorization header.
Does LM Studio process multiple requests in parallel?+
Yes, since version 0.4.0, which introduced parallel requests to the same model with continuous batching. The Max Concurrent Predictions setting determines the number of simultaneous requests; beyond that, requests wait. For intensive multi-user use, with dozens of simultaneous requests, a server designed for this purpose, such as vLLM, remains more suitable.
What identifier should you put in the model field?+
The model identifier as it appears in LM Studio, not a generic name. The /v1/models request lists them. With on-demand loading enabled, it returns all downloaded models; without it, only those already loaded into memory. Copy the exact identifier into your client, because a typo causes the request to fail.
Is LM Studio free for business use?+
Yes: since July 8, 2025, LM Studio has been free at home and at work, with no commercial license to request, according to the publisher's announcement. Commercial offerings exist for additional needs. Check the current terms of use before deployment, especially if you plan to use the publisher's commercial offerings.
Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.