What does Mixture-of-experts serving mean in inference serving?
Mixture-of-experts serving routes each token through a selected subset of available experts. Total parameter storage can greatly exceed the parameters active for one token. Operators place experts across devices, set routing capacity, and observe token distribution under real prompts. They measure load skew, communication, cache use, and latency by expert path.