Concepts

Prefill and decode

Reference entryHow LLMs work
Reference entryPrefill and decode

Prefill processes all prompt tokens in parallel to initialize model state; decode then generates new tokens sequentially using that cached state.

Where it sits