What does Paged attention mean in inference serving?
Paged attention stores each sequence's KV cache in fixed-size blocks rather than one contiguous allocation. A block table maps logical token positions to physical cache blocks during attention. The runtime allocates blocks as sequences grow, then returns blocks when sequences finish or caches expire. Block size and eviction policy trade metadata work against wasted memory and cache reuse.