Open source

vLLM

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

vLLM is open-source software that runs large language models on a server and answers many requests at once, using GPU memory carefully so fewer bytes go to waste.

1 · What it is

vLLM is open-source software for serving large language models, the programs behind chatbots. Serving means running a model on a server so many apps and people can ask it questions at the same time. vLLM began at the University of California, Berkeley, is now built by a large community, and is released under the Apache 2.0 licence.

Its best-known idea is PagedAttention. While a model writes an answer, it keeps working notes called the KV cache, short for key-value cache, on the graphics card. Those notes grow as the answer gets longer. The vLLM team found that older servers wasted 60 to 80 percent of this memory, because they reserved big chunks in advance. PagedAttention borrows a trick from operating systems, the software like Windows or Linux that runs a computer. Think of a car park attendant who hands out one space at a time instead of holding a whole row. PagedAttention cuts the notes into small, equal blocks, hands out a new block only when one is needed, and tracks where each block lives in a table. Waste drops to under 4 percent. That means many more requests fit in one batch, a group of requests the graphics card works on together. In the original paper, vLLM handled two to four times more requests per second than earlier systems, without making replies slower.

You start a vLLM server with one command, such as vllm serve followed by a model name. It copies the OpenAI API, the set of rules apps use to talk to OpenAI’s models. So apps already written for that service can point at your own server instead. vLLM also adds new requests to a batch as they arrive. It can reuse work when prompts start the same way. It runs on NVIDIA, AMD and Intel graphics cards as well as ordinary CPUs.

2 · Why it exists

Serving a language model to many people at once runs into three snags.

Memory runs outEach request keeps working notes on the graphics card, and those notes grow as the reply gets longer.
Space gets wastedOlder systems reserved big chunks of memory up front, and much of it sat empty or in awkward gaps.
Apps need a familiar doorPrograms want to send questions the same way they already talk to hosted AI services.
3 · How it works

Follow several chat requests through one vLLM server.

How vLLM serves many requests with paged KV cache memory Five boxes in a row. Several chat requests arrive through the OpenAI-compatible API on port 8000. The scheduler batches them. The highlighted step, PagedAttention, splits each request's KV cache into fixed-size blocks and records them in a block table. The blocks can sit anywhere in GPU memory and are added only as new tokens appear. Answers go back to the client app. A dashed path shows that requests with the same start can share blocks. MANY REQUESTS, MEMORY HANDED OUT IN SMALL BLOCKS Chat requests Scheduler PagedAttention GPU memory Replies sent to the OpenAI-compatible API on port 8000 batches new requests as they arrive splits KV cache into fixed blocks, notes a block table blocks sit anywhere, added only when needed answers go back to the client app same start: shared
The highlighted box is PagedAttention: working memory is handed out in small blocks, like pages, instead of one big reserved chunk.
  1. 1 · launchYou start the server with one command naming a model, and it listens on port 8000 by default.
  2. 2 · sendApps send prompts through routes that copy the OpenAI API.
  3. 3 · batchThe server batches incoming requests together so the graphics card stays busy.
  4. 4 · pagePagedAttention stores each request's KV cache in small blocks and adds a new block only when needed.
  5. 5 · replyThe model generates each answer and sends it back to the app.

Less wasted memory means more requests fit on the same graphics card at once.

4 · Where it's used
WhoWhat they askWhat it works with
Startup team“Can our existing OpenAI client code talk to an open model we host ourselves?”The OpenAI-compatible server started with vllm serve
Chatbot team“Can we avoid recomputing the same long system prompt for every user?”Automatic prefix caching
Research lab“How many chat requests can one graphics card handle at the same time?”PagedAttention and continuous batching
Platform engineer“Will it run on our AMD GPUs or plain CPUs?”vLLM's hardware support list
5 · What it solves, and what it doesn't
solves
  • It serves language models to many users with high throughput.
  • It cuts wasted KV cache memory so more requests fit in one batch.
  • It offers an OpenAI-compatible server so existing client code can connect.
  • It can reuse work for prompts that share the same start.
doesn't solve
  • It is built for running models, not training them.
  • The basic server hosts one model at a time.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. repovllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs, vLLM project · read 28 Sept 2026
  2. repovLLM LICENSE, vLLM project · read 28 Sept 2026
  3. paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv (SOSP 2023) · read 28 Sept 2026
  4. officialvLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention, vLLM Blog · read 28 Sept 2026
  5. docsQuickstart, vLLM Documentation · read 28 Sept 2026
  6. docsAutomatic Prefix Caching, vLLM Documentation · read 28 Sept 2026