What does Disaggregated serving mean in inference serving?
Disaggregated serving places prefill and decode work on separate worker groups. KV cache state moves between the groups when a request changes phases. Routing, cache format, network topology, and backpressure determine whether separation helps. Systems must recover when transfer or the destination worker fails after prefill.