RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Continuous batching

Batching

Continuous batching

What it is

Continuous batching schedules work at each model iteration instead of waiting for an entire request batch to finish. Finished sequences leave the active batch while ready sequences enter without restarting the serving loop.

Why it moves cost and latency

This keeps accelerator slots useful across requests with different prompt and output lengths. Higher utilization can lower cost per token, but crowded batches can increase queueing and tail latency.

What it looks like in practice

A serving scheduler admits new sequences between decode steps and removes sequences that have completed. Operators tune batch limits against memory pressure, latency targets, and workload shape.

Where we measured it

  • Measured serving receipts under load->

Related terms

  • Paged attention->
  • KV cache->
  • TTFT->
  • Inter-token latency->
  • Tokens per second->
  • Concurrency->
  • Chunked prefill->
  • Continuous versus static batching->
  • In-flight batching->

Questions this definition answers

What does Continuous batching mean in inference serving?

Continuous batching schedules work at each model iteration instead of waiting for an entire request batch to finish. Finished sequences leave the active batch while ready sequences enter without restarting the serving loop. A serving scheduler admits new sequences between decode steps and removes sequences that have completed. Operators tune batch limits against memory pressure, latency targets, and workload shape.

Why can Continuous batching move cost or latency?

This keeps accelerator slots useful across requests with different prompt and output lengths. Higher utilization can lower cost per token, but crowded batches can increase queueing and tail latency.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy