RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Tensor versus pipeline parallelism

Parallelism

Tensor versus pipeline parallelism

What it is

Tensor parallelism splits work within model layers and exchanges partial results during those layers. Pipeline parallelism places whole layer ranges on stages and communicates activations at stage boundaries.

Why it moves cost and latency

Tensor sharding favors fast repeated collectives, while pipeline staging depends on balanced work and enough microbatches to limit bubbles. Model size, topology, batch shape, and latency targets decide the fit.

What it looks like in practice

Teams compare feasible layouts using the same model, devices, and request distribution. Memory headroom, collective time, stage imbalance, and end-to-end latency remain visible.

Related terms

  • Tensor parallelism->
  • Pipeline parallelism->
  • NVLink->
  • HBM->

Questions this definition answers

What does Tensor versus pipeline parallelism mean in inference serving?

Tensor parallelism splits work within model layers and exchanges partial results during those layers. Pipeline parallelism places whole layer ranges on stages and communicates activations at stage boundaries. Teams compare feasible layouts using the same model, devices, and request distribution. Memory headroom, collective time, stage imbalance, and end-to-end latency remain visible.

Why can Tensor versus pipeline parallelism move cost or latency?

Tensor sharding favors fast repeated collectives, while pipeline staging depends on balanced work and enough microbatches to limit bubbles. Model size, topology, batch shape, and latency targets decide the fit.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy