RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Batching

Glossary category

Batching glossary

Definitions for inference serving terms in the batching category.

Terms in this category

  • Continuous batchingContinuous batching schedules work at each model iteration instead of waiting for an entire request batch to finish.
  • ConcurrencyConcurrency is the number of client requests in flight within a stated measurement boundary.
  • Chunked prefillChunked prefill divides a long prompt's prefill work into smaller token blocks.
  • Continuous versus static batchingStatic batching fixes a request group until every sequence in that group finishes.
  • In-flight batchingIn-flight batching is a common name for continuous batching in serving systems.

Related glossary categories

  • Memory glossary->
  • Quantization glossary->
  • Latency glossary->
  • Parallelism glossary->
  • Serving glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy