What does CUDA graphs mean in inference serving?
CUDA graphs record a repeatable sequence of device operations and replay it with fewer host launches. The captured graph expects stable execution structure and memory addresses. Serving systems may capture common shape buckets, pad work, or recapture when shapes change. Unsupported paths remain eager, and graph memory must stay valid across replays.