Open source

SGLang

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

SGLang is open-source server software that runs large language models on GPUs and answers many requests quickly, partly by reusing work shared between prompts.

1 · What it is

SGLang is open-source software for serving large language models, the programs behind chatbots. Serving means running a model on a server so many apps and people can send it questions at once. SGLang aims to answer quickly and to handle lots of requests. It can run on one graphics card (GPU) or on a large group of them working together. It is released under the Apache 2.0 licence.

Its best-known trick is called RadixAttention. When a model reads a prompt, it builds working notes called the KV cache, short for key-value cache. Instead of throwing those notes away after a request, SGLang keeps them in a radix tree, a kind of index sorted by the words each prompt starts with. When a new prompt begins the same way, for example with an earlier chat turn, SGLang reuses the saved notes and only processes the new words. When memory fills up, it removes the notes that were used least recently.

Apps talk to SGLang the same way they talk to OpenAI’s API, the doorway programs use to send questions to a model. So code written for OpenAI can point at your own server instead. SGLang can also make replies follow a fixed format such as JSON, a tidy data layout that other programs can read. It works with most models published on Hugging Face and runs on NVIDIA and AMD GPUs, Google TPUs and other hardware.

2 · Why it exists

Serving a language model to many people at once has three snags.

Repeated workDifferent prompts can share the same start, such as an earlier chat turn, and working through it again wastes memory and computation.
Limited GPU memoryGPU memory is limited, so a server cannot keep every saved note and must choose what to throw away.
Apps need a standard doorPrograms already written for OpenAI's API need a server that speaks the same way.
3 · How it works

Follow two chat requests that begin with the same words.

How SGLang reuses the shared start of two prompts Five boxes in a row. Two requests arrive that begin with the same system prompt. The SGLang server receives the full prompts. The highlighted step, RadixAttention, finds the shared start in a radix tree of stored KV cache and reuses it. The GPU computes only the new words. Replies return through an OpenAI-compatible API. Least recently used entries are evicted when memory fills. TWO REQUESTS, ONE SHARED START, COMPUTED ONCE Two requests SGLang server RadixAttention GPU computes Replies same start, then a different question gets the full prompts on a local port finds the shared start in a radix tree, reuses KV only the new words of each prompt sent back via the OpenAI- compatible API new KV cache kept When GPU memory fills, the least recently used entries in the tree are evicted.
The highlighted box is RadixAttention: the shared start of a prompt is looked up and reused instead of computed again.
  1. 1 · launchYou start the SGLang server with a model, and it waits for requests on a local port.
  2. 2 · sendAn app sends full prompts, for example through SGLang's OpenAI-compatible API.
  3. 3 · matchRadixAttention searches a radix tree for any stored start of the prompt and reuses its KV cache.
  4. 4 · computeThe GPU only works through the new part of each prompt and generates the reply.
  5. 5 · keepThe new KV cache is kept in the tree, and the least recently used entries are removed when memory runs short.

The app does nothing special: the server finds and reuses shared prompt starts on its own.

4 · Where it's used
WhoWhat they askWhat it works with
Chatbot team“Can we avoid recomputing the same earlier chat turn for every new message?”Prefix caching with RadixAttention
App developer“Can my existing OpenAI client code talk to a model we host ourselves?”The OpenAI-compatible chat completions route on the SGLang server
Data pipeline builder“Can the model be forced to return valid JSON?”Structured outputs with JSON, regex or EBNF constraints
Research lab“Can the same engine run on NVIDIA, AMD or TPU hardware?”SGLang's hardware support list
5 · What it solves, and what it doesn't
solves
  • It serves language models, including multimodal ones, quickly and at high volume.
  • It reuses the KV cache for prompts that share a start, without manual setup.
  • It offers OpenAI-compatible APIs so existing client code can connect.
  • It can constrain output to JSON, regex or EBNF formats.
doesn't solve
  • Reuse only helps when requests actually share the same starting text.
  • It does not train models; it serves them.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. reposgl-project/sglang: SGLang is a high-performance serving framework for large language models and multimodal models, SGLang project · read 28 Sept 2026
  2. repoSGLang LICENSE, SGLang project · read 28 Sept 2026
  3. paperSGLang: Efficient Execution of Structured Language Model Programs, arXiv · read 28 Sept 2026
  4. officialFast and Expressive LLM Inference with RadixAttention and SGLang, LMSYS Org · read 28 Sept 2026
  5. docsSending Requests, SGLang Documentation · read 28 Sept 2026
  6. docsOpenAI APIs - Completions, SGLang Documentation · read 28 Sept 2026