Concepts

Throughput

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Throughput is the amount of completed inference work divided by the time spent measuring it.

1 · What it is

Request throughput counts final responses per benchmark duration. Output-token throughput counts generated output tokens over that duration.

Batching can raise throughput by processing several requests together. Waiting to form batches can also add latency. MLPerf requires benchmark latency constraints to be met for a successful test.

Under batching, one model execution can serve more than one inference.

2 · Why it exists

A serving system can measure completed inference work per second.

Count the unitOutput-token throughput counts output tokens over benchmark duration.
Define completionNVIDIA defines request throughput as final responses divided by benchmark duration.
Load changes resultsIncreasing batch size may improve throughput while increasing latency.
3 · How it works

Count completed requests and output tokens inside one measurement window.

State both the completed-work unit and the measurement duration.
  1. 1 · loadSpecify the model, server URL and input type for the benchmark.
  2. 2 · serveQueue requests and buffer them to form a batch.
  3. 3 · countCount output tokens completed during the window.
  4. 4 · divideDivide the completed count by benchmark duration.

Request throughput counts final responses per benchmark duration.

4 · Where it's used
WhoWhat they askWhat it works with
Capacity planner“How many requests can this deployment finish each second?”Final responses per benchmark duration
LLM serving engineer“How many output tokens does the fleet produce each second?”Aggregate output-token throughput
Performance engineer“Did a larger batch improve rate at an acceptable wait time?”Throughput-latency measurements across load levels
5 · What it solves, and what it doesn't
solves
  • It quantifies completed work per unit time.
  • MLCommons uses tokens per second to measure LLM throughput.
doesn't solve
  • High throughput does not guarantee low per-request latency.
  • The amount of computation in one LLM inference can differ from another.
  • Counting model executions is not the same as counting individual inferences under batching.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsGenAI-Perf metrics, NVIDIA · read 28 Sept 2026
  2. docsModel Analyzer metrics, NVIDIA · read 28 Sept 2026
  3. docsTriton Statistics Extension, NVIDIA · read 28 Sept 2026
  4. docsServe a Text Generator with Request Batching, Ray · read 28 Sept 2026
  5. officialLlama 2 70B - An MLPerf Inference Benchmark for Large Language Models, MLCommons · read 28 Sept 2026