Efficient Memory Management for Large Language Model Serving with PagedAttention
Date: 2nd October 2026
Related: ELI5: FlashAttention
Key Points
- KV cache is veruy inefficient in contiguous, static memory as prompts can be different lengths. Old methods just pre-assign maximum length, which is highly wasteful.
- vLLM introduces paged memory and paged attention:
- Paged Memory: a scedhuler and system for storing data in non-contiguous chunks. Saves memory by not having to reserve huge chunks.
- Paged Attention: an efficient kernel for computing attention across chunms of the Qs, Ks and Vs
- Bunch of inference methods