What does Tensor versus pipeline parallelism mean in inference serving?
Tensor parallelism splits work within model layers and exchanges partial results during those layers. Pipeline parallelism places whole layer ranges on stages and communicates activations at stage boundaries. Teams compare feasible layouts using the same model, devices, and request distribution. Memory headroom, collective time, stage imbalance, and end-to-end latency remain visible.