RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Grouped-query attention

Memory

Grouped-query attention

What it is

Grouped-query attention assigns several query heads to each shared key-value head. It stores fewer distinct key and value streams than full multi-head attention while retaining multiple query heads.

Why it moves cost and latency

Fewer key-value heads shrink cache footprint and bandwidth requirements during generation. The model architecture chooses the sharing ratio, so serving software must follow the checkpoint design.

What it looks like in practice

Attention kernels map each query-head group to its corresponding key-value head. Operators verify kernel support, cache layout, and quality for the model they deploy.

Related terms

  • KV cache->
  • Multi-query attention->
  • HBM->

Questions this definition answers

What does Grouped-query attention mean in inference serving?

Grouped-query attention assigns several query heads to each shared key-value head. It stores fewer distinct key and value streams than full multi-head attention while retaining multiple query heads. Attention kernels map each query-head group to its corresponding key-value head. Operators verify kernel support, cache layout, and quality for the model they deploy.

Why can Grouped-query attention move cost or latency?

Fewer key-value heads shrink cache footprint and bandwidth requirements during generation. The model architecture chooses the sharing ratio, so serving software must follow the checkpoint design.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy