RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Flash attention

Memory

Flash attention

What it is

Flash attention is an IO-aware exact attention method that tiles computation around limited on-chip memory. It avoids materializing the full attention score matrix in high-bandwidth memory.

Why it moves cost and latency

Fewer transfers between device memory and on-chip storage reduce attention memory traffic. The benefit depends on sequence shape, supported kernels, and the surrounding model workload.

What it looks like in practice

A compatible kernel processes query, key, and value tiles while maintaining numerically stable running statistics. Kernel selection and tile sizes vary with data type, head shape, and accelerator.

Related terms

  • Paged attention->
  • Prefill->
  • Kernel fusion->
  • HBM->

Questions this definition answers

What does Flash attention mean in inference serving?

Flash attention is an IO-aware exact attention method that tiles computation around limited on-chip memory. It avoids materializing the full attention score matrix in high-bandwidth memory. A compatible kernel processes query, key, and value tiles while maintaining numerically stable running statistics. Kernel selection and tile sizes vary with data type, head shape, and accelerator.

Why can Flash attention move cost or latency?

Fewer transfers between device memory and on-chip storage reduce attention memory traffic. The benefit depends on sequence shape, supported kernels, and the surrounding model workload.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy