Latency
Request latency is the elapsed time between sending an inference request and receiving its final response.
Latency needs two boundaries. For a streaming model, time to first token measures from request send to the first streamed output. Request latency continues until the final response.
Queue duration isolates time spent waiting for scheduling. Inter-token latency records gaps between consecutive streamed outputs after streaming begins.
GenAI-Perf reports average, p90 and p99 latency statistics. Metric terminology is not standardized across benchmarking tools. MLPerf’s server scenario models online applications with random query arrival.
Request latency ends when the final response is received.
Mark one request from send time through its first and final responses.
- 1 · startRecord when the benchmark sends the request.
- 2 · queueMeasure time spent waiting for the scheduler.
- 3 · firstStop the TTFT interval when the first response arrives.
- 4 · streamMeasure gaps between consecutive streamed outputs.
- 5 · finishStop request latency when the final response arrives.
Metric terminology is not standardized across benchmarking tools.
| Who | What they ask | What it works with |
|---|---|---|
| Product engineer | “How long before a user sees the first token?” | Time to first token |
| Serving operator | “Are requests waiting before model execution?” | Queue duration |
| Reliability engineer | “How slow are the worst successful requests?” | Tail-latency percentiles |
- Request latency provides one value per request in a benchmark.
- TTFT measures from request send to the first streamed output.
- Queue duration measures time spent waiting for scheduling.
- GenAI-Perf reports averages and percentile latency statistics.
- Large batch sizes can add a latency penalty for the first requests.
- Metric terminology is not standardized across benchmarking tools.
- MLPerf's server scenario models online applications with random query arrival.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGenAI-Perf metrics, NVIDIA · read 28 Sept 2026
- docsTriton Inference Server metrics, NVIDIA · read 28 Sept 2026
- docsvLLM benchmark CLI, vLLM · read 28 Sept 2026
- docsDynamic Request Batching, Ray · read 28 Sept 2026
- officialLlama 2 70B - An MLPerf Inference Benchmark for Large Language Models, MLCommons · read 28 Sept 2026