RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Concurrency

Batching

Concurrency

What it is

Concurrency is the number of client requests in flight within a stated measurement boundary. It can include queued and executing work, so it is not necessarily the active model batch size.

Why it moves cost and latency

More concurrent work can improve accelerator utilization and aggregate throughput until memory or compute saturates. Beyond that point, queueing and batch contention can raise per-request latency and tail spread.

What it looks like in practice

Benchmarks hold prompt and output shapes steady while sweeping in-flight requests at the client boundary. Production systems set admission limits from cache capacity, latency objectives, and expected traffic bursts.

Where we measured it

  • Benchmark receipts with disclosed load->
  • Throughput and latency by concurrency->
  • Serving-engine concurrency sweep->
  • Measured concurrency comparison rows->

Related terms

  • Continuous batching->
  • Paged attention->
  • KV cache->
  • Prefix caching->
  • TTFT->
  • Inter-token latency->
  • Tokens per second->
  • p50 vs p99->
  • Tensor parallelism->
  • Context length->
  • Admission control->
  • Data-parallel serving->

Questions this definition answers

What does Concurrency mean in inference serving?

Concurrency is the number of client requests in flight within a stated measurement boundary. It can include queued and executing work, so it is not necessarily the active model batch size. Benchmarks hold prompt and output shapes steady while sweeping in-flight requests at the client boundary. Production systems set admission limits from cache capacity, latency objectives, and expected traffic bursts.

Why can Concurrency move cost or latency?

More concurrent work can improve accelerator utilization and aggregate throughput until memory or compute saturates. Beyond that point, queueing and batch contention can raise per-request latency and tail spread.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy