RunInfraby RightNow
  • CatalogNew
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Tensor parallelism

Parallelism

Tensor parallelism

What it is

Tensor parallelism shards weight matrices and their computation across multiple accelerators within each model layer. Participating devices exchange partial results so the layer produces the same logical output.

Why it moves cost and latency

Sharding lets a model use combined memory and compute, but communication occurs repeatedly through the network fabric. Slow collectives or small workloads can leave communication cost larger than the local compute saved.

What it looks like in practice

A serving deployment chooses a parallel group size that fits the model and the available interconnect. Operators compare memory headroom and latency because adding devices does not guarantee a faster request.

Related terms

  • FP8 quantization->
  • Inter-token latency->
  • Tokens per second->
  • Concurrency->
  • Pipeline parallelism->
  • Tensor versus pipeline parallelism->
  • Expert parallelism->
  • NVLink->

Questions this definition answers

What does Tensor parallelism mean in inference serving?

Tensor parallelism shards weight matrices and their computation across multiple accelerators within each model layer. Participating devices exchange partial results so the layer produces the same logical output. A serving deployment chooses a parallel group size that fits the model and the available interconnect. Operators compare memory headroom and latency because adding devices does not guarantee a faster request.

Why can Tensor parallelism move cost or latency?

Sharding lets a model use combined memory and compute, but communication occurs repeatedly through the network fabric. Slow collectives or small workloads can leave communication cost larger than the local compute saved.

If you need custom optimization for a specific model, describe what you need

Describe the model and hardware you want optimized...
ModelsAuto engineAuto GPU
End-to-end encryption
Isolated GPU infrastructure
No training on your data
SOC 2 Type II
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModelsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy