RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Paged attention

Memory

Paged attention

What it is

Paged attention stores each sequence's KV cache in fixed-size blocks rather than one contiguous allocation. A block table maps logical token positions to physical cache blocks during attention.

Why it moves cost and latency

Block allocation reduces memory fragmentation and lets the server reuse freed capacity between sequences. Better cache occupancy can support larger active batches before memory becomes the limiting resource.

What it looks like in practice

The runtime allocates blocks as sequences grow, then returns blocks when sequences finish or caches expire. Block size and eviction policy trade metadata work against wasted memory and cache reuse.

Related terms

  • Continuous batching->
  • KV cache->
  • Prefix caching->
  • Concurrency->
  • Flash attention->
  • Sliding-window attention->

Questions this definition answers

What does Paged attention mean in inference serving?

Paged attention stores each sequence's KV cache in fixed-size blocks rather than one contiguous allocation. A block table maps logical token positions to physical cache blocks during attention. The runtime allocates blocks as sequences grow, then returns blocks when sequences finish or caches expire. Block size and eviction policy trade metadata work against wasted memory and cache reuse.

Why can Paged attention move cost or latency?

Block allocation reduces memory fragmentation and lets the server reuse freed capacity between sequences. Better cache occupancy can support larger active batches before memory becomes the limiting resource.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy