RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Inter-token latency

Latency

Inter-token latency

What it is

Inter-token latency is the elapsed time between consecutive generated tokens after the first token. It tracks the cadence of decode steps experienced by a request.

Why it moves cost and latency

Slower decode cadence extends response completion time even when the first token arrives quickly. Batch contention, cache access, kernel efficiency, and cross-device communication all influence the interval.

What it looks like in practice

A serving trace timestamps each streamed token and computes intervals after generation begins. Teams inspect both typical and tail intervals because pauses can hide behind a good average.

Where we measured it

  • Measured inter-token latency receipts->

Related terms

  • Continuous batching->
  • KV cache->
  • Tokens per second->
  • p50 vs p99->
  • Concurrency->
  • Speculative decoding->
  • Tensor parallelism->
  • Decode->
  • CUDA graphs->
  • Streaming responses->
  • Guided decoding->
  • Sampling parameters->

Questions this definition answers

What does Inter-token latency mean in inference serving?

Inter-token latency is the elapsed time between consecutive generated tokens after the first token. It tracks the cadence of decode steps experienced by a request. A serving trace timestamps each streamed token and computes intervals after generation begins. Teams inspect both typical and tail intervals because pauses can hide behind a good average.

Why can Inter-token latency move cost or latency?

Slower decode cadence extends response completion time even when the first token arrives quickly. Batch contention, cache access, kernel efficiency, and cross-device communication all influence the interval.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy