RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Memory

Glossary category

Memory glossary

Definitions for inference serving terms in the memory category.

Terms in this category

  • Paged attentionPaged attention stores each sequence's KV cache in fixed-size blocks rather than one contiguous allocation.
  • KV cacheThe KV cache stores attention keys and values computed for earlier tokens in each active sequence.
  • Prefix cachingPrefix caching preserves KV cache blocks for prompt prefixes that recur across requests.
  • Flash attentionFlash attention is an IO-aware exact attention method that tiles computation around limited on-chip memory.
  • Grouped-query attentionGrouped-query attention assigns several query heads to each shared key-value head.
  • Multi-query attentionMulti-query attention gives all query heads one shared key head and one shared value head.
  • Sliding-window attentionSliding-window attention limits each token to keys and values within a recent fixed-size window.
  • Attention sinksAttention sinks are retained early tokens that receive attention after later tokens move through a sliding window.
  • RoPE scalingRoPE scaling changes how rotary position frequencies or positions are mapped during inference.
  • Context lengthContext length is the maximum number of token positions a serving configuration admits for a sequence.

Related glossary categories

  • Batching glossary->
  • Quantization glossary->
  • Latency glossary->
  • Parallelism glossary->
  • Serving glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy