RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/TTFT

Latency

TTFT

What it is

Time to first token measures elapsed time from request arrival until the first generated token becomes available. It includes queueing and setup, while prompt prefill usually supplies most active model computation.

Why it moves cost and latency

Long first-token delay makes interactive systems feel unresponsive before generation has visibly started. Prefix reuse, prompt length, queueing, and prefill throughput can move this metric independently of decode speed.

What it looks like in practice

Teams record the request boundary and the first streamed token, then inspect the distribution across workloads. They separate queue delay from prefill when diagnosing whether capacity or prompt processing is responsible.

Where we measured it

  • First-token latency by concurrency->
  • First-token latency across serving engines->
  • First-token latency comparison rows->

Related terms

  • Continuous batching->
  • Prefix caching->
  • p50 vs p99->
  • Concurrency->
  • RoPE scaling->
  • Prefill->
  • Disaggregated serving->
  • Request queueing->
  • Model loading->

Questions this definition answers

What does TTFT mean in inference serving?

Time to first token measures elapsed time from request arrival until the first generated token becomes available. It includes queueing and setup, while prompt prefill usually supplies most active model computation. Teams record the request boundary and the first streamed token, then inspect the distribution across workloads. They separate queue delay from prefill when diagnosing whether capacity or prompt processing is responsible.

Why can TTFT move cost or latency?

Long first-token delay makes interactive systems feel unresponsive before generation has visibly started. Prefix reuse, prompt length, queueing, and prefill throughput can move this metric independently of decode speed.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy