Concepts

PagedAttention

Reference entryInference and serving
Reference entryPagedAttention

PagedAttention stores a model's attention memory in small fixed-size blocks, like pages in computer memory, so a server can fit and share more requests at once; it was introduced with vLLM.

Where it sits