Concepts
Prefill and decode
Reference entryPrefill and decode
Prefill processes all prompt tokens in parallel to initialize model state; decode then generates new tokens sequentially using that cached state.
Where it sits
Prefill processes all prompt tokens in parallel to initialize model state; decode then generates new tokens sequentially using that cached state.