Concepts
Tokens per second
1 · In one line
Tokens per second reports how many tokens a server returns in a set amount of time.
1 · What it is
Tokens per second reports how many tokens a server returns in a set amount of time. Text Generation Inference exposes generated tokens per request.
Throughput and latency are orthogonal measurements. Serving systems can compare throughput at the same latency level.
KV-cache management can limit batch size and server throughput. Batching can improve overall throughput.
Serving tools can count generated tokens for each request.
Request counterText Generation Inference exposes generated tokens per request.
Latency is separateThroughput and latency are orthogonal measurements.
Serving phasesDistServe names TTFT for prefill and TPOT for decoding.
Workload changes resultsBatching more requests can improve overall throughput.
Keep request counters, server throughput and latency distinct.
- 1 · requestText Generation Inference exposes generated tokens per request.
- 2 · serveThe server returns tokens over a set amount of time.
- 3 · reportReport that server throughput as a token rate.
- 4 · compareCompare latency separately from throughput.
Throughput and latency are orthogonal measurements.
| Who | What they ask | What it works with |
|---|---|---|
| Model engineer | “How quickly does one request stream after its first token?” | Request output tokens per second |
| Serving operator | “How much generation work does the deployment finish?” | Aggregate output tokens per second |
| Capacity planner | “Did a serving change raise throughput at the same latency target?” | Throughput under a latency constraint |
solves
- It reports how many tokens a server returns over time.
- Text Generation Inference exposes generated tokens per request.
- Serving systems can compare throughput at the same latency level.
doesn't solve
- Throughput and latency are orthogonal measurements.
- KV-cache management can limit batch size and server throughput.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsBenchmarking Hugging Face Text Generation Inference, Hugging Face · read 28 Sept 2026
- docsText Generation Inference metrics, Hugging Face · read 28 Sept 2026
- paperDistServe - Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving, arXiv · read 28 Sept 2026
- paperEfficient Memory Management for Large Language Model Serving with PagedAttention, ACM SOSP / arXiv · read 28 Sept 2026
- paperSarathi-Serve - Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve, arXiv · read 28 Sept 2026