RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Serving

Glossary category

Serving glossary

Definitions for inference serving terms in the serving category.

Terms in this category

  • Disaggregated servingDisaggregated serving places prefill and decode work on separate worker groups.
  • torch.compiletorch.
  • Admission controlAdmission control decides which queued requests enter active model execution and which remain waiting.
  • Request queueingRequest queueing is the wait between server arrival and admission to model execution.
  • Streaming responsesStreaming responses send generated tokens to the caller as they become available, commonly through server-sent events.
  • Offline batch inferenceOffline batch inference processes stored inputs without an interactive caller waiting for each response.
  • Cold startA cold start is the delay before an inactive or new replica can accept useful model work.
  • Model loadingModel loading reads checkpoint tensors and places the required weights into host or device memory.
  • SafeTensorsSafeTensors is a tensor checkpoint format with a metadata header and raw tensor data.
  • GGUFGGUF is a self-describing checkpoint format used widely by CPU and edge-oriented runtimes.
  • Mixture-of-experts servingMixture-of-experts serving routes each token through a selected subset of available experts.
  • LoRA servingLoRA serving applies low-rank adapter updates to a shared base model during inference.
  • Guided decodingGuided decoding constrains generation so emitted text follows a grammar, schema, or other formal rule.

Related glossary categories

  • Batching glossary->
  • Memory glossary->
  • Quantization glossary->
  • Latency glossary->
  • Parallelism glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy