What does Chunked prefill mean in inference serving?
Chunked prefill divides a long prompt's prefill work into smaller token blocks. A scheduler can place those blocks beside decode iterations instead of running the whole prompt uninterrupted. Operators tune chunk limits with batch capacity, prompt distribution, and latency objectives. Implementations differ in how chunks share an iteration with decode work.