RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Prefill

Latency

Prefill

What it is

Prefill is the parallel model pass over prompt tokens that creates their KV cache state. On accelerators, sufficiently large prefills are often limited by compute throughput.

Why it moves cost and latency

Prefill supplies much of the active model work before the first generated token. Prompt length, batching, prefix reuse, and queueing therefore shape time to first token.

What it looks like in practice

Schedulers batch prompt tokens together or divide long prompts into chunks around decode work. Teams separate prefill duration from queue delay when diagnosing first-token latency.

Where we measured it

  • Measured first-token latency receipts->
  • First-token latency by serving workload->

Related terms

  • TTFT->
  • Flash attention->
  • Chunked prefill->
  • Decode->
  • Disaggregated serving->

Questions this definition answers

What does Prefill mean in inference serving?

Prefill is the parallel model pass over prompt tokens that creates their KV cache state. On accelerators, sufficiently large prefills are often limited by compute throughput. Schedulers batch prompt tokens together or divide long prompts into chunks around decode work. Teams separate prefill duration from queue delay when diagnosing first-token latency.

Why can Prefill move cost or latency?

Prefill supplies much of the active model work before the first generated token. Prompt length, batching, prefix reuse, and queueing therefore shape time to first token.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy