RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Decode

Latency

Decode

What it is

Decode is the sequential generation phase after prefill, with each iteration producing one token per active sequence. It repeatedly reads model weights and growing cache state.

Why it moves cost and latency

For common serving shapes, decode is often limited by memory bandwidth rather than arithmetic throughput. Its iteration cadence drives inter-token latency and response completion time.

What it looks like in practice

A scheduler advances active sequences together while removing completions and admitting eligible work. Teams inspect token intervals, batch occupancy, cache traffic, and kernel launch overhead.

Where we measured it

  • Measured decode cadence and throughput->

Related terms

  • Inter-token latency->
  • Prefill->
  • Disaggregated serving->
  • CUDA graphs->
  • Kernel fusion->
  • HBM->

Questions this definition answers

What does Decode mean in inference serving?

Decode is the sequential generation phase after prefill, with each iteration producing one token per active sequence. It repeatedly reads model weights and growing cache state. A scheduler advances active sequences together while removing completions and admitting eligible work. Teams inspect token intervals, batch occupancy, cache traffic, and kernel launch overhead.

Why can Decode move cost or latency?

For common serving shapes, decode is often limited by memory bandwidth rather than arithmetic throughput. Its iteration cadence drives inter-token latency and response completion time.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy