What does Expert parallelism mean in inference serving?
Expert parallelism places mixture-of-experts feed-forward experts on different devices. Each token is routed to selected experts, then expert outputs return to the token's execution path. Deployments choose expert placement, routing capacity, and communication groups around the available topology. Operators track load skew, dropped or rerouted work, and collective time.