RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Tokens per second

Latency

Tokens per second

What it is

Tokens per second expresses how quickly a serving system produces tokens over an interval. It may describe one request's generation rate or aggregate throughput across many active requests.

Why it moves cost and latency

Aggregate throughput determines how much demand a deployment can serve for a given amount of accelerator time. Per-request rate describes user-visible generation speed, so the two forms should never be mixed.

What it looks like in practice

Benchmarks state whether they count output tokens, total tokens, or per-request decode rate. They pair the rate with concurrency, prompt shape, output shape, and latency so the result remains interpretable.

Where we measured it

  • Measured output throughput receipts->
  • Serving-engine throughput sweep->

Related terms

  • Continuous batching->
  • FP8 quantization->
  • Inter-token latency->
  • p50 vs p99->
  • Concurrency->
  • Speculative decoding->
  • Tensor parallelism->
  • Kernel fusion->
  • torch.compile->
  • Offline batch inference->
  • Mixture-of-experts serving->
  • LoRA serving->

Questions this definition answers

What does Tokens per second mean in inference serving?

Tokens per second expresses how quickly a serving system produces tokens over an interval. It may describe one request's generation rate or aggregate throughput across many active requests. Benchmarks state whether they count output tokens, total tokens, or per-request decode rate. They pair the rate with concurrency, prompt shape, output shape, and latency so the result remains interpretable.

Why can Tokens per second move cost or latency?

Aggregate throughput determines how much demand a deployment can serve for a given amount of accelerator time. Per-request rate describes user-visible generation speed, so the two forms should never be mixed.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy