Concepts
Throughput
1 · In one line
Throughput is the amount of completed inference work divided by the time spent measuring it.
1 · What it is
Request throughput counts final responses per benchmark duration. Output-token throughput counts generated output tokens over that duration.
Batching can raise throughput by processing several requests together. Waiting to form batches can also add latency. MLPerf requires benchmark latency constraints to be met for a successful test.
Under batching, one model execution can serve more than one inference.
A serving system can measure completed inference work per second.
Count the unitOutput-token throughput counts output tokens over benchmark duration.
Define completionNVIDIA defines request throughput as final responses divided by benchmark duration.
Load changes resultsIncreasing batch size may improve throughput while increasing latency.
Count completed requests and output tokens inside one measurement window.
- 1 · loadSpecify the model, server URL and input type for the benchmark.
- 2 · serveQueue requests and buffer them to form a batch.
- 3 · countCount output tokens completed during the window.
- 4 · divideDivide the completed count by benchmark duration.
Request throughput counts final responses per benchmark duration.
| Who | What they ask | What it works with |
|---|---|---|
| Capacity planner | “How many requests can this deployment finish each second?” | Final responses per benchmark duration |
| LLM serving engineer | “How many output tokens does the fleet produce each second?” | Aggregate output-token throughput |
| Performance engineer | “Did a larger batch improve rate at an acceptable wait time?” | Throughput-latency measurements across load levels |
solves
- It quantifies completed work per unit time.
- MLCommons uses tokens per second to measure LLM throughput.
doesn't solve
- High throughput does not guarantee low per-request latency.
- The amount of computation in one LLM inference can differ from another.
- Counting model executions is not the same as counting individual inferences under batching.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGenAI-Perf metrics, NVIDIA · read 28 Sept 2026
- docsModel Analyzer metrics, NVIDIA · read 28 Sept 2026
- docsTriton Statistics Extension, NVIDIA · read 28 Sept 2026
- docsServe a Text Generator with Request Batching, Ray · read 28 Sept 2026
- officialLlama 2 70B - An MLPerf Inference Benchmark for Large Language Models, MLCommons · read 28 Sept 2026