vLLM
vLLM is open-source software that runs large language models on a server and answers many requests at once, using GPU memory carefully so fewer bytes go to waste.
vLLM is open-source software for serving large language models, the programs behind chatbots. Serving means running a model on a server so many apps and people can ask it questions at the same time. vLLM began at the University of California, Berkeley, is now built by a large community, and is released under the Apache 2.0 licence.
Its best-known idea is PagedAttention. While a model writes an answer, it keeps working notes called the KV cache, short for key-value cache, on the graphics card. Those notes grow as the answer gets longer. The vLLM team found that older servers wasted 60 to 80 percent of this memory, because they reserved big chunks in advance. PagedAttention borrows a trick from operating systems, the software like Windows or Linux that runs a computer. Think of a car park attendant who hands out one space at a time instead of holding a whole row. PagedAttention cuts the notes into small, equal blocks, hands out a new block only when one is needed, and tracks where each block lives in a table. Waste drops to under 4 percent. That means many more requests fit in one batch, a group of requests the graphics card works on together. In the original paper, vLLM handled two to four times more requests per second than earlier systems, without making replies slower.
You start a vLLM server with one command, such as vllm serve followed by a model name. It copies the OpenAI API, the set of rules apps use to talk to OpenAI’s models. So apps already written for that service can point at your own server instead. vLLM also adds new requests to a batch as they arrive. It can reuse work when prompts start the same way. It runs on NVIDIA, AMD and Intel graphics cards as well as ordinary CPUs.
Serving a language model to many people at once runs into three snags.
Follow several chat requests through one vLLM server.
- 1 · launchYou start the server with one command naming a model, and it listens on port 8000 by default.
- 2 · sendApps send prompts through routes that copy the OpenAI API.
- 3 · batchThe server batches incoming requests together so the graphics card stays busy.
- 4 · pagePagedAttention stores each request's KV cache in small blocks and adds a new block only when needed.
- 5 · replyThe model generates each answer and sends it back to the app.
Less wasted memory means more requests fit on the same graphics card at once.
| Who | What they ask | What it works with |
|---|---|---|
| Startup team | “Can our existing OpenAI client code talk to an open model we host ourselves?” | The OpenAI-compatible server started with vllm serve |
| Chatbot team | “Can we avoid recomputing the same long system prompt for every user?” | Automatic prefix caching |
| Research lab | “How many chat requests can one graphics card handle at the same time?” | PagedAttention and continuous batching |
| Platform engineer | “Will it run on our AMD GPUs or plain CPUs?” | vLLM's hardware support list |
- It serves language models to many users with high throughput.
- It cuts wasted KV cache memory so more requests fit in one batch.
- It offers an OpenAI-compatible server so existing client code can connect.
- It can reuse work for prompts that share the same start.
- It is built for running models, not training them.
- The basic server hosts one model at a time.
Sources used
This explainer is written in original language. The links below support its factual claims.
- repovllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs, vLLM project · read 28 Sept 2026
- repovLLM LICENSE, vLLM project · read 28 Sept 2026
- paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv (SOSP 2023) · read 28 Sept 2026
- officialvLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention, vLLM Blog · read 28 Sept 2026
- docsQuickstart, vLLM Documentation · read 28 Sept 2026
- docsAutomatic Prefix Caching, vLLM Documentation · read 28 Sept 2026