SGLang: serve a local LLM to multiple users utilisateurs
SGLang is an inference server designed for multiple simultaneous users: it installs with uv, starts with a single command on port 30000, and gets its throughput under load from RadixAttention (prefix caching) and continuous request scheduling. It requires a recent CUDA GPU, making it a shared-server tool, not a replacement for Ollama on a single-user personal workstation.
SGLang is an inference server framework developed by the LMSYS community, designed for high-throughput, low-latency serving across multiple simultaneous requests. This guide covers its installation, how RadixAttention works, the concurrency controls you need to know, and the criteria for choosing between SGLang, vLLM, and Ollama based on the actual number of users you need to serve.
#What SGLang does
SGLang presents itself as a high-performance serving framework for LLMs and multimodal models, designed for low-latency, high-throughput inference, from a single GPU to large distributed clusters. The project claims production deployments generating trillions of tokens every day on more than 400,000 GPUs worldwide, and is hosted by the nonprofit open-source organization LMSYS.
Compatible with the OpenAI and Hugging Face APIs, SGLang supports a wide range of models (Llama, Qwen, DeepSeek, GLM, Mistral, Gemma) and hardware (GPU NVIDIA, AMD, Intel Xeon CPU, Google TPU, Ascend NPU). The latest version as of September 28, 2026 is v0.5.20, released on September 18, 2026.
#Install and start the server
Deploy local AI at work: privacy, compliance, multi-user architecture, costs, the one-page memo for leadership.
- Lifetime online access
- PDF + files
- Lifetime updates
The installation recommended by the official documentation uses uv, which is faster than standard pip. The --prerelease=allow flag is required because some SGLang dependencies publish only pre-releases on PyPI.
Launching in a Docker container remains the most reproducible approach for a server deployment, with port 30000 used by default for the API.
The documentation states that SGLang now requires CUDA 13: CUDA 12 images and wheels (cu129) have been removed because PyTorch 2.14 no longer publishes a CUDA 12.9 build, and version 0.5.19 remains the last to offer a CUDA 12 path. A deployment on older hardware must therefore pin this version or verify the GPU driver before updating.
#RadixAttention and prefix caching
The core of SGLang’s throughput promise is RadixAttention: already-computed sequence prefixes (a shared system prompt, the beginning of a multi-turn conversation) are organized in a radix tree and reused across requests instead of being recomputed on every call. The project’s original announcement, in January 2024, claimed up to 5x faster inference thanks to this mechanism—a project-announcement figure from its own benchmark at the time, not a recent independent measurement on your hardware.
This improvement mainly benefits scenarios with shared prefixes: multiple users querying the same system prompt, an agent rereading the same context on every turn, or few-shot prompting with the same examples. A request stream with no shared prefix at all (completely independent questions with no history) benefits much less from RadixAttention.
The runtime combines RadixAttention with a zero-overhead CPU scheduler, prefill-decode disaggregation, speculative decoding, continuous request scheduling (continuous batching), and paged attention—a stack of optimization techniques rather than a single mechanism.
#Tune concurrency and memory
Three launch parameters govern how SGLang handles multiple users in parallel. --mem-fraction-static sets the share of GPU memory reserved for the model weights and KV cache; the documentation recommends reducing it when you get an out-of-memory error, otherwise it is calculated automatically from the available GPU memory.
| Parameter | Role |
|---|---|
| --max-running-requests | Maximum number of requests processed simultaneously (no default limit) |
| --max-queued-requests | Maximum number of requests queued before processing |
| --schedule-policy | Request scheduling policy: fcfs (first come, first served) by default, or lpm, random, dfs-weight, lof, priority, routing-key |
| --chunked-prefill-size | Splitting prefill into chunks to prevent a long request from blocking the others |
Without an explicit limit on --max-running-requests, SGLang accepts as many requests as the KV cache memory allows, which can degrade per-request latency under heavy load instead of politely refusing new connections. Set an explicit limit consistent with the available VRAM before opening the server to multiple real users.
#When to choose SGLang over vLLM or Ollama
The three tools address different needs. Ollama targets single-user personal use with minimal setup; SGLang and vLLM target multi-user service on server GPUs, with similar approaches (prefix caching, continuous batching) but distinct histories and ecosystems.
- Single user, personal workstation
- Ollama or llama.cpp remain easier to install and run on consumer CPUs or GPUs, without network-service configuration.
- Multiple users, with a system already built around vLLM
- Sticking with vLLM avoids a migration; our dedicated guide covers deploying it in production.
- Multiple users, with throughput prioritized under shared prefixes
- SGLang, with RadixAttention, is the natural candidate—provided you have a compatible CUDA GPU.
- Desired Ollama client compatibility without installing Ollama
- SGLang exposes an API compatible with the CLI and Python library Ollama, allowing you to reuse existing scripts without running the Ollama server itself.
For a broader overview of inference architectures (llama.cpp, vLLM, Exllama), the site's backend comparison details the tradeoffs beyond the SGLang use case alone.
#What the hardware dictates
SGLang's quick-start guide is explicit: a NVIDIA GPU with CUDA sm80 or higher support (A10, A100, L4, L40S, H100) is a prerequisite for the standard Linux installation path, the recommended platform. The project also announces broader hardware support — AMD GPUs (MI355, MI300), Intel Xeon CPUs, Google TPUs, and Ascend NPUs — through dedicated installation paths distinct from the main NVIDIA GPU path.
#What memory costs per request
“Serving multiple users” very concretely means VRAM consumption grows with the number of active requests, not just with model size. The KV cache managed by --mem-fraction-static and --max-total-tokens stores two vectors (key and value) per attention layer for each token already generated by a request. The general formula is: bytes per token = 2 × number of layers × KV attention heads × head dimension × bytes per value (2 in FP16/BF16, 1 in FP8).
Illustrative calculation for an architecture close to Llama-3.1-8B-Instruct (32 layers, 8 grouped-query attention KV heads, head dimension 128): in FP16, this yields 2 × 32 × 8 × 128 × 2 = 131,072 bytes per token, or approximately 128 KB per token. For a request with 8,192 context tokens (prompt and response combined), the KV cache for that request alone then uses approximately 1 GB of VRAM—before even counting the model weights. On a 24 GB GPU with approximately 16 GB reserved for the Q4 weights and runtime base, the remaining headroom only allows a few simultaneous active requests at this context length, which explains why an explicit limit on --max-running-requests prevents degradation instead of cleanly rejecting new connections.
#Security and observability
By default, the SGLang API launched with sglang.launch_server requires no authentication: anyone who can reach port 30000 can send requests. The --api-key option defines a key required by the server, including on its OpenAI-compatible endpoint; without it, exposing the server beyond localhost or a trusted internal network amounts to leaving inference access open, along with the associated GPU bill.
For production monitoring, the --enable-metrics option (disabled by default) publishes Prometheus-format metrics on a /metrics endpoint: request throughput, inference latency, token generation speed, and cache efficiency are exposed for regular scraping, rather than inferring server status from logs alone.
- Security principles for an exposed inference server
- Source: server argument reference (--api-key, --enable-metrics)
#Troubleshooting: symptoms, cause, fix
| Symptom | Likely cause | Correction |
|---|---|---|
| Out-of-memory error at startup | --mem-fraction-static calculated automatically too high for the VRAM actually available | Explicitly reduce --mem-fraction-static, as recommended by the official documentation |
| Latency spikes under load without errors | No limit on --max-running-requests: the server accepts requests as long as the KV cache allows | Set --max-running-requests based on the available VRAM and the per-request cost calculation above |
| Installation or startup failure after an update | Upgrade to a version that requires CUDA 13 while the driver remains on CUDA 12 | Pin version 0.5.19 (the last CUDA 12 path) or update the GPU driver before SGLang |
| Server accessible from the network without controls | No key defined: --api-key is empty by default | Set --api-key before exposing the server beyond a trusted network |
#Limitations and points to watch
The “up to 5x faster” figure that still accompanies the RadixAttention presentation dates back to the project’s initial announcement in January 2024: it illustrates the mechanism’s benefit on the benchmark used at the time, not a guaranteed gain on a recent deployment with different models and GPUs. Actual sizing depends on the prefix-sharing rate between requests, the model selected, and the GPU used—it should be measured on your own traffic rather than inferred from the original announcement.
The mandatory move to CUDA 13 (with version 0.5.19 being the last to offer a CUDA 12 path) warrants checking the GPU driver and installed CUDA version before upgrading to a recent version, or installation may fail on a server still running CUDA 12.
The project's release cadence also needs monitoring for production deployments: SGLang releases minor versions roughly every two to three weeks (v0.5.20 on September 18, 2026, v0.5.19 on September 5, v0.5.18 on August 22), which means pinning a specific version in Docker images instead of following the latest tag on a production service, at the risk of an unexpected behavior change during redeployment.
- Deploy vLLM in production
- vLLM: what is it, who is it for, and when should you use it?
- llama.cpp vs vLLM vs Exllama
- Source: official SGLang documentation
- Source: official README of the SGLang repository
- Source: server argument reference
Does SGLang replace Ollama?+
Is a GPU required to run SGLang?+
Does RadixAttention speed up all requests in the same way?+
Which parameter should you tune first to serve multiple users with SGLang?+
Does SGLang work with an older GPU limited to CUDA 12?+
How much GPU memory does the KV cache for a single request use?+
How do you secure an SGLang server exposed on the network?+
Feedback, an error, or a clarification? Let us know—it improves the guide for everyone.