RunInfraby RightNow
  • Pricing
  • Research
  • Contact
DashboardSign inGet started
Home/Glossary/Parallelism

Glossary category

Parallelism glossary

Definitions for inference serving terms in the parallelism category.

Terms in this category

  • Tensor parallelismTensor parallelism shards weight matrices and their computation across multiple accelerators within each model layer.
  • Pipeline parallelismPipeline parallelism divides consecutive model layers into stages placed on different devices.
  • Tensor versus pipeline parallelismTensor parallelism splits work within model layers and exchanges partial results during those layers.
  • Expert parallelismExpert parallelism places mixture-of-experts feed-forward experts on different devices.
  • Data-parallel servingData-parallel serving runs independent model replicas and routes different requests to each replica.

Related glossary categories

  • Batching glossary->
  • Memory glossary->
  • Quantization glossary->
  • Latency glossary->
  • Serving glossary->
  • Hardware glossary->
RunInfraby RightNow

© 2026 RunInfra. All rights reserved.

System status
Pipeline BuilderModel APIsCost CalculatorPricingStartupsBenchmarksDocsResearchNewsContact
Backed by
YCombinator
AICPA Type II
SOC 2
NVIDIA Inception ProgramNVIDIA Inception Program
Ask AI about RunInfra
Part of RightNow
SecurityDPAAUPCookiesTermsPrivacy