Advanced 12 minDeployment

SGLang: serve a local LLM to multiple users utilisateurs

Direct response

SGLang is an inference server designed for multiple simultaneous users: it installs with uv, starts with a single command on port 30000, and gets its throughput under load from RadixAttention (prefix caching) and continuous request scheduling. It requires a recent CUDA GPU, making it a shared-server tool, not a replacement for Ollama on a single-user personal workstation.

SGLang is an inference server framework developed by the LMSYS community, designed for high-throughput, low-latency serving across multiple simultaneous requests. This guide covers its installation, how RadixAttention works, the concurrency controls you need to know, and the criteria for choosing between SGLang, vLLM, and Ollama based on the actual number of users you need to serve.

By Mohamed Meguedmi·Update 2026-09-28·Tested on Windows, macOS, and Linux

#What SGLang does

SGLang presents itself as a high-performance serving framework for LLMs and multimodal models, designed for low-latency, high-throughput inference, from a single GPU to large distributed clusters. The project claims production deployments generating trillions of tokens every day on more than 400,000 GPUs worldwide, and is hosted by the nonprofit open-source organization LMSYS.

Compatible with the OpenAI and Hugging Face APIs, SGLang supports a wide range of models (Llama, Qwen, DeepSeek, GLM, Mistral, Gemma) and hardware (GPU NVIDIA, AMD, Intel Xeon CPU, Google TPU, Ascend NPU). The latest version as of September 28, 2026 is v0.5.20, released on September 18, 2026.

i
SGLang is not a direct replacement for Ollama
SGLang targets multi-user serving on server hardware. It even exposes an API compatible with the Ollama client to simplify migration from existing tools, but it neither installs nor replaces Ollama itself.

#Install and start the server

The AI at Work Kit

Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.

  • Lifetime online access
  • PDF + files
  • Lifetime updates

The installation recommended by the official documentation uses uv, which is faster than standard pip. The --prerelease=allow flag is required because some SGLang dependencies publish only pre-releases on PyPI.

Installation
pip install --upgrade pip
pip install uv
uv pip install --prerelease=allow sglang

Launching in a Docker container remains the most reproducible approach for a server deployment, with port 30000 used by default for the API.

Docker
docker run --gpus all \
    --shm-size 32g \
    -p 30000:30000 \
    -v ~/.cache/huggingface:/root/.cache/huggingface \
    --env "HF_TOKEN=<secret>" \
    --ipc=host \
    lmsysorg/sglang:latest \
    python3 -m sglang.launch_server --model-path meta-llama/Llama-3.1-8B-Instruct --host 0.0.0.0 --port 30000

The documentation states that SGLang now requires CUDA 13: CUDA 12 images and wheels (cu129) have been removed because PyTorch 2.14 no longer publishes a CUDA 12.9 build, and version 0.5.19 remains the last to offer a CUDA 12 path. A deployment on older hardware must therefore pin this version or verify the GPU driver before updating.

#RadixAttention and prefix caching

The core of SGLang’s throughput promise is RadixAttention: already-computed sequence prefixes (a shared system prompt, the beginning of a multi-turn conversation) are organized in a radix tree and reused across requests instead of being recomputed on every call. The project’s original announcement, in January 2024, claimed up to 5x faster inference thanks to this mechanism—a project-announcement figure from its own benchmark at the time, not a recent independent measurement on your hardware.

This improvement mainly benefits scenarios with shared prefixes: multiple users querying the same system prompt, an agent rereading the same context on every turn, or few-shot prompting with the same examples. A request stream with no shared prefix at all (completely independent questions with no history) benefits much less from RadixAttention.

→
Surprise gradient
RadixAttention's benefit depends directly on the proportion of prefix tokens shared between requests. A deployment where each user has their own long, different system prompt loses much of the cache's value, even though the feature remains active.

The runtime combines RadixAttention with a zero-overhead CPU scheduler, prefill-decode disaggregation, speculative decoding, continuous request scheduling (continuous batching), and paged attention—a stack of optimization techniques rather than a single mechanism.

#Tune concurrency and memory

Three launch parameters govern how SGLang handles multiple users in parallel. --mem-fraction-static sets the share of GPU memory reserved for the model weights and KV cache; the documentation recommends reducing it when you get an out-of-memory error, otherwise it is calculated automatically from the available GPU memory.

Concurrency parameters to know
ParameterRole
--max-running-requestsMaximum number of requests processed simultaneously (no default limit)
--max-queued-requestsMaximum number of requests queued before processing
--schedule-policyRequest scheduling policy: fcfs (first come, first served) by default, or lpm, random, dfs-weight, lof, priority, routing-key
--chunked-prefill-sizeSplitting prefill into chunks to prevent a long request from blocking the others

Without an explicit limit on --max-running-requests, SGLang accepts as many requests as the KV cache memory allows, which can degrade per-request latency under heavy load instead of politely refusing new connections. Set an explicit limit consistent with the available VRAM before opening the server to multiple real users.

#When to choose SGLang over vLLM or Ollama

The three tools address different needs. Ollama targets single-user personal use with minimal setup; SGLang and vLLM target multi-user service on server GPUs, with similar approaches (prefix caching, continuous batching) but distinct histories and ecosystems.

Single user, personal workstation
Ollama or llama.cpp remain easier to install and run on consumer CPUs or GPUs, without network-service configuration.
Multiple users, with a system already built around vLLM
Sticking with vLLM avoids a migration; our dedicated guide covers deploying it in production.
Multiple users, with throughput prioritized under shared prefixes
SGLang, with RadixAttention, is the natural candidate—provided you have a compatible CUDA GPU.
Desired Ollama client compatibility without installing Ollama
SGLang exposes an API compatible with the CLI and Python library Ollama, allowing you to reuse existing scripts without running the Ollama server itself.

For a broader overview of inference architectures (llama.cpp, vLLM, Exllama), the site's backend comparison details the tradeoffs beyond the SGLang use case alone.

#What the hardware dictates

SGLang's quick-start guide is explicit: a NVIDIA GPU with CUDA sm80 or higher support (A10, A100, L4, L40S, H100) is a prerequisite for the standard Linux installation path, the recommended platform. The project also announces broader hardware support — AMD GPUs (MI355, MI300), Intel Xeon CPUs, Google TPUs, and Ascend NPUs — through dedicated installation paths distinct from the main NVIDIA GPU path.

!
Not a CPU-only tool by default
Unlike llama.cpp or Ollama, SGLang is not primarily designed to run without a dedicated GPU. On a Mac or a PC without a recent NVIDIA card, Ollama or LM Studio remain the default choices; SGLang makes the most sense on a dedicated GPU server serving multiple users.

#What memory costs per request

“Serving multiple users” very concretely means VRAM consumption grows with the number of active requests, not just with model size. The KV cache managed by --mem-fraction-static and --max-total-tokens stores two vectors (key and value) per attention layer for each token already generated by a request. The general formula is: bytes per token = 2 × number of layers × KV attention heads × head dimension × bytes per value (2 in FP16/BF16, 1 in FP8).

Illustrative calculation for an architecture close to Llama-3.1-8B-Instruct (32 layers, 8 grouped-query attention KV heads, head dimension 128): in FP16, this yields 2 × 32 × 8 × 128 × 2 = 131,072 bytes per token, or approximately 128 KB per token. For a request with 8,192 context tokens (prompt and response combined), the KV cache for that request alone then uses approximately 1 GB of VRAM—before even counting the model weights. On a 24 GB GPU with approximately 16 GB reserved for the Q4 weights and runtime base, the remaining headroom only allows a few simultaneous active requests at this context length, which explains why an explicit limit on --max-running-requests prevents degradation instead of cleanly rejecting new connections.

→
Order of magnitude, not a measurement
This calculation is a theoretical estimate based on the model architecture, not a figure measured on a real SGLang deployment: the exact number of simultaneous requests supported also depends on the loaded model, the actual context length, and --mem-fraction-static.

#Security and observability

By default, the SGLang API launched with sglang.launch_server requires no authentication: anyone who can reach port 30000 can send requests. The --api-key option defines a key required by the server, including on its OpenAI-compatible endpoint; without it, exposing the server beyond localhost or a trusted internal network amounts to leaving inference access open, along with the associated GPU bill.

For production monitoring, the --enable-metrics option (disabled by default) publishes Prometheus-format metrics on a /metrics endpoint: request throughput, inference latency, token generation speed, and cache efficiency are exposed for regular scraping, rather than inferring server status from logs alone.

#Troubleshooting: symptoms, cause, fix

Common symptoms in a multi-user service
SymptomLikely causeCorrection
Out-of-memory error at startup--mem-fraction-static calculated automatically too high for the VRAM actually availableExplicitly reduce --mem-fraction-static, as recommended by the official documentation
Latency spikes under load without errorsNo limit on --max-running-requests: the server accepts requests as long as the KV cache allowsSet --max-running-requests based on the available VRAM and the per-request cost calculation above
Installation or startup failure after an updateUpgrade to a version that requires CUDA 13 while the driver remains on CUDA 12Pin version 0.5.19 (the last CUDA 12 path) or update the GPU driver before SGLang
Server accessible from the network without controlsNo key defined: --api-key is empty by defaultSet --api-key before exposing the server beyond a trusted network

#Limitations and points to watch

The “up to 5x faster” figure that still accompanies the RadixAttention presentation dates back to the project’s initial announcement in January 2024: it illustrates the mechanism’s benefit on the benchmark used at the time, not a guaranteed gain on a recent deployment with different models and GPUs. Actual sizing depends on the prefix-sharing rate between requests, the model selected, and the GPU used—it should be measured on your own traffic rather than inferred from the original announcement.

The mandatory move to CUDA 13 (with version 0.5.19 being the last to offer a CUDA 12 path) warrants checking the GPU driver and installed CUDA version before upgrading to a recent version, or installation may fail on a server still running CUDA 12.

The project's release cadence also needs monitoring for production deployments: SGLang releases minor versions roughly every two to three weeks (v0.5.20 on September 18, 2026, v0.5.19 on September 5, v0.5.18 on August 22), which means pinning a specific version in Docker images instead of following the latest tag on a production service, at the risk of an unexpected behavior change during redeployment.

Frequently asked questions
Does SGLang replace Ollama?+
No, they're two different tools. Ollama targets single-user personal use on a workstation, with minimal setup; SGLang targets service for multiple simultaneous users on a server GPU, with more advanced configuration. SGLang even exposes an API compatible with the Ollama client, without Ollama actually running behind it—useful for reusing existing scripts without migrating the entire toolchain.
Is a GPU required to run SGLang?+
The official quick-start guide requires a NVIDIA GPU with CUDA sm80 support or higher (A10, A100, L4, L40S, H100) for the standard Linux path. Separate paths exist for AMD, Intel Xeon, TPU, or NPU, but SGLang isn't primarily designed for CPU-only use on a personal computer: Ollama or llama.cpp remain better suited in that case.
Does RadixAttention speed up all requests in the same way?+
No. The benefit depends on the volume of shared prefix tokens across requests: a system prompt shared by multiple users or a conversation history reused on every turn benefits greatly. Completely independent requests, with no common prefix between them, benefit much less, even though the mechanism remains active on the server.
Which parameter should you tune first to serve multiple users with SGLang?+
--max-running-requests, to set an explicit limit consistent with the available VRAM and the KV-cache cost of each active request. Without this limit, SGLang accepts requests as long as the KV cache allows, which can degrade per-request latency under heavy load instead of politely rejecting additional connections.
Does SGLang work with an older GPU limited to CUDA 12?+
Not with recent versions. SGLang now requires CUDA 13; version 0.5.19 is the last to offer a CUDA 12 path: a server still running on a CUDA 12 driver must either pin this version or update its GPU driver before installing a newer version of SGLang, or the installation will fail.
How much GPU memory does the KV cache for a single request use?+
It depends on the model and context length, but you can estimate the order of magnitude: for an architecture close to Llama-3.1-8B (32 layers, 8 KV heads), a context of 8,192 tokens uses about 1 GB of VRAM in FP16, before even counting the model weights. Switching to FP8 cuts that figure in half.
How do you secure an SGLang server exposed on the network?+
Set --api-key to require token authentication for all requests, including the OpenAI-compatible endpoint; without this option, the server remains open to anyone who can reach port 30000. Enable --enable-metrics separately to monitor throughput and latency through Prometheus, without confusing observability with access control.

Did this guide help you?

Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.