RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Latency

Glossary category

Latency glossary

Definitions for inference serving terms in the latency category.

Terms in this category

  • TTFTTime to first token measures elapsed time from request arrival until the first generated token becomes available.
  • Inter-token latencyInter-token latency is the elapsed time between consecutive generated tokens after the first token.
  • Tokens per secondTokens per second expresses how quickly a serving system produces tokens over an interval.
  • p50 vs p99The fiftieth and ninety-ninth percentiles mark different positions in an ordered latency distribution.
  • Speculative decodingSpeculative decoding uses a cheaper draft process to propose several candidate tokens before the target model verifies them.
  • PrefillPrefill is the parallel model pass over prompt tokens that creates their KV cache state.
  • DecodeDecode is the sequential generation phase after prefill, with each iteration producing one token per active sequence.
  • CUDA graphsCUDA graphs record a repeatable sequence of device operations and replay it with fewer host launches.
  • Kernel fusionKernel fusion combines adjacent tensor operations into a single device kernel.
  • Sampling parametersSampling parameters reshape the next-token distribution after model logits are computed.

Related glossary categories

  • Batching glossary->
  • Memory glossary->
  • Quantization glossary->
  • Parallelism glossary->
  • Serving glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy