SGLang
SGLang is open-source server software that runs large language models on GPUs and answers many requests quickly, partly by reusing work shared between prompts.
SGLang is open-source software for serving large language models, the programs behind chatbots. Serving means running a model on a server so many apps and people can send it questions at once. SGLang aims to answer quickly and to handle lots of requests. It can run on one graphics card (GPU) or on a large group of them working together. It is released under the Apache 2.0 licence.
Its best-known trick is called RadixAttention. When a model reads a prompt, it builds working notes called the KV cache, short for key-value cache. Instead of throwing those notes away after a request, SGLang keeps them in a radix tree, a kind of index sorted by the words each prompt starts with. When a new prompt begins the same way, for example with an earlier chat turn, SGLang reuses the saved notes and only processes the new words. When memory fills up, it removes the notes that were used least recently.
Apps talk to SGLang the same way they talk to OpenAI’s API, the doorway programs use to send questions to a model. So code written for OpenAI can point at your own server instead. SGLang can also make replies follow a fixed format such as JSON, a tidy data layout that other programs can read. It works with most models published on Hugging Face and runs on NVIDIA and AMD GPUs, Google TPUs and other hardware.
Serving a language model to many people at once has three snags.
Follow two chat requests that begin with the same words.
- 1 · launchYou start the SGLang server with a model, and it waits for requests on a local port.
- 2 · sendAn app sends full prompts, for example through SGLang's OpenAI-compatible API.
- 3 · matchRadixAttention searches a radix tree for any stored start of the prompt and reuses its KV cache.
- 4 · computeThe GPU only works through the new part of each prompt and generates the reply.
- 5 · keepThe new KV cache is kept in the tree, and the least recently used entries are removed when memory runs short.
The app does nothing special: the server finds and reuses shared prompt starts on its own.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot team | “Can we avoid recomputing the same earlier chat turn for every new message?” | Prefix caching with RadixAttention |
| App developer | “Can my existing OpenAI client code talk to a model we host ourselves?” | The OpenAI-compatible chat completions route on the SGLang server |
| Data pipeline builder | “Can the model be forced to return valid JSON?” | Structured outputs with JSON, regex or EBNF constraints |
| Research lab | “Can the same engine run on NVIDIA, AMD or TPU hardware?” | SGLang's hardware support list |
- It serves language models, including multimodal ones, quickly and at high volume.
- It reuses the KV cache for prompts that share a start, without manual setup.
- It offers OpenAI-compatible APIs so existing client code can connect.
- It can constrain output to JSON, regex or EBNF formats.
- Reuse only helps when requests actually share the same starting text.
- It does not train models; it serves them.
Sources used
This explainer is written in original language. The links below support its factual claims.
- reposgl-project/sglang: SGLang is a high-performance serving framework for large language models and multimodal models, SGLang project · read 28 Sept 2026
- repoSGLang LICENSE, SGLang project · read 28 Sept 2026
- paperSGLang: Efficient Execution of Structured Language Model Programs, arXiv · read 28 Sept 2026
- officialFast and Expressive LLM Inference with RadixAttention and SGLang, LMSYS Org · read 28 Sept 2026
- docsSending Requests, SGLang Documentation · read 28 Sept 2026
- docsOpenAI APIs - Completions, SGLang Documentation · read 28 Sept 2026