Concepts

Tokens per second

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Tokens per second reports how many tokens a server returns in a set amount of time.

1 · What it is

Tokens per second reports how many tokens a server returns in a set amount of time. Text Generation Inference exposes generated tokens per request.

Throughput and latency are orthogonal measurements. Serving systems can compare throughput at the same latency level.

KV-cache management can limit batch size and server throughput. Batching can improve overall throughput.

2 · Why it exists

Serving tools can count generated tokens for each request.

Request counterText Generation Inference exposes generated tokens per request.
Latency is separateThroughput and latency are orthogonal measurements.
Serving phasesDistServe names TTFT for prefill and TPOT for decoding.
Workload changes resultsBatching more requests can improve overall throughput.
3 · How it works

Keep request counters, server throughput and latency distinct.

Serving systems expose generated-token counters, throughput and latency as distinct measurements.
  1. 1 · requestText Generation Inference exposes generated tokens per request.
  2. 2 · serveThe server returns tokens over a set amount of time.
  3. 3 · reportReport that server throughput as a token rate.
  4. 4 · compareCompare latency separately from throughput.

Throughput and latency are orthogonal measurements.

4 · Where it's used
WhoWhat they askWhat it works with
Model engineer“How quickly does one request stream after its first token?”Request output tokens per second
Serving operator“How much generation work does the deployment finish?”Aggregate output tokens per second
Capacity planner“Did a serving change raise throughput at the same latency target?”Throughput under a latency constraint
5 · What it solves, and what it doesn't
solves
  • It reports how many tokens a server returns over time.
  • Text Generation Inference exposes generated tokens per request.
  • Serving systems can compare throughput at the same latency level.
doesn't solve
  • Throughput and latency are orthogonal measurements.
  • KV-cache management can limit batch size and server throughput.