RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/HBM

Hardware

HBM

What it is

High-bandwidth memory is stacked device memory placed close to an accelerator's compute package. It offers much more bandwidth than typical host memory but remains finite in capacity.

Why it moves cost and latency

Decode throughput often depends on moving weights and cache state through this memory quickly. Compute rate alone cannot describe serving performance when memory traffic is the limiting work.

What it looks like in practice

Operators budget capacity among weights, KV cache, workspaces, and captured graphs. Precision, batch shape, access pattern, and model architecture determine bandwidth pressure.

Related terms

  • KV cache->
  • Flash attention->
  • Grouped-query attention->
  • Multi-query attention->
  • Context length->
  • Decode->
  • Tensor versus pipeline parallelism->
  • Mixture-of-experts serving->

Questions this definition answers

What does HBM mean in inference serving?

High-bandwidth memory is stacked device memory placed close to an accelerator's compute package. It offers much more bandwidth than typical host memory but remains finite in capacity. Operators budget capacity among weights, KV cache, workspaces, and captured graphs. Precision, batch shape, access pattern, and model architecture determine bandwidth pressure.

Why can HBM move cost or latency?

Decode throughput often depends on moving weights and cache state through this memory quickly. Compute rate alone cannot describe serving performance when memory traffic is the limiting work.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy