RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Model loading

Serving

Model loading

What it is

Model loading reads checkpoint tensors and places the required weights into host or device memory. Checkpoint layout, shard count, storage locality, and conversion work shape the path.

Why it moves cost and latency

Loading time contributes to cold starts, recovery, and replica scaling. Extra copies or format conversion can consume memory headroom before the model serves a request.

What it looks like in practice

Loaders validate metadata, stream or map tensor data, and allocate destination buffers. Readiness begins only after required transfers, initialization, and warmup have completed.

Related terms

  • TTFT->
  • Cold start->
  • SafeTensors->
  • GGUF->
  • Offline batch inference->
  • Mixture-of-experts serving->
  • LoRA serving->

Questions this definition answers

What does Model loading mean in inference serving?

Model loading reads checkpoint tensors and places the required weights into host or device memory. Checkpoint layout, shard count, storage locality, and conversion work shape the path. Loaders validate metadata, stream or map tensor data, and allocate destination buffers. Readiness begins only after required transfers, initialization, and warmup have completed.

Why can Model loading move cost or latency?

Loading time contributes to cold starts, recovery, and replica scaling. Extra copies or format conversion can consume memory headroom before the model serves a request.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy