Building with AI

Model serving

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Model serving means running a trained model on a server so apps can send it requests over the network and get answers back.

1 · What it is

Training a model and using it are two different jobs. Model serving is the work of loading a trained model onto a server, a computer that waits for requests, and giving it a network address. Any app can then send the model a request and get an answer back, the same way it would call a weather or maps service.

A homework-help app, for example, sends a student’s question to an address such as /v1/chat/completions. vLLM, an open-source engine for serving large language models, provides a server that answers at addresses like that. It accepts requests in the same format as OpenAI’s Completions API. So any HTTP client, meaning any program that can send web requests, can talk to it.

TensorFlow Serving does a similar job for TensorFlow models. It is built for production, meaning real apps with real users. It is designed so a team can add new algorithms and run experiments without changing the server design or its APIs.

NVIDIA’s Triton Inference Server is another open-source serving system. It can load models built with many tools, including PyTorch, ONNX and TensorRT. Apps can reach it in two ways: HTTP/REST, the usual web style, or gRPC, another way for programs to send requests. Triton can also run several models at the same time on one server, which it calls concurrent model execution.

The hard part is answering many requests at once, quickly, without wasting expensive hardware. Three ideas do most of the work.

Batching

Triton has a dynamic batcher. It groups requests on the server into one batch. That usually raises throughput, meaning answers per second. You can let the batcher wait a short, set time so more requests can join. That wait is the catch. Triton’s own guide describes it as trading more latency, the delay each user feels, for more throughput. By default its batcher does not wait at all. The guide suggests raising the wait step by step and stopping once the delay goes past your latency budget, the longest wait you are willing to accept.

Picture the homework app on a school night. Three students press “send” within a moment of each other. With no wait, the server may run each question separately. With a short wait, all three questions travel through the model together as one batch. Each student waits slightly longer, but the server finishes more answers per second.

Memory for language models

Language models add a twist. Each request keeps a working memory called the KV cache, short for key-value cache. The vLLM research paper explains that this memory is large for each request and grows and shrinks as the answer is written. If it is managed badly, memory is wasted, and that waste limits how many requests fit in one batch. The paper’s fix, PagedAttention, borrows a memory trick called paging from operating systems. The authors report 2 to 4 times more throughput than earlier systems at the same latency. vLLM also uses a method called continuous batching for incoming requests.

Scaling

The third idea is scaling, meaning running more or fewer copies of the model as demand changes. KServe builds on Kubernetes, a system that runs software across many machines, and adds settings made for running models. It handles load balancing (spreading requests across copies), autoscaling (adding or removing copies automatically), canary rollouts and monitoring. A canary rollout lets a team test a new model version on part of the traffic. KServe’s autoscaler can watch how fast the model is writing answers (token throughput, where a token is a small piece of a word), how long the waiting line gets and how busy the GPUs (the chips that run the model) are. It can shrink to zero copies when nobody is calling. It can also cope with sudden rushes of traffic.

Watching and safety

Serving also means watching. Triton reports metrics such as GPU use, server throughput and server latency, so a team can see when answers slow down. Security needs care too. vLLM’s own docs warn not to rely on its API key option alone to protect the server.

What serving does not do is make the model smarter. It runs whatever model it is given, as fast and as cheaply as it can.

2 · Why it exists

A trained model does nothing until something runs it and lets apps reach it.

No front doorAn app needs a network address it can send requests to, such as an HTTP endpoint.
One at a time is slowGrouping requests into batches usually lets a server answer more requests per second.
Traffic changesDemand rises and falls, so the number of running copies has to change with it.
3 · How it works

Follow requests from several apps through one model server.

The batcher trades a small wait for more answers per second.
  1. 1 · loadThe server loads the trained model and opens an HTTP endpoint.
  2. 2 · queueRequests from many apps arrive and wait in a queue.
  3. 3 · batchThe batcher combines waiting requests into one batch, waiting no longer than a set limit.
  4. 4 · runThe model runs the batch and each answer goes back to its app.
  5. 5 · scaleAn autoscaler adds or removes model copies as traffic changes.

Batching raises throughput (answers per second) but can add latency (wait per answer).

4 · Where it's used
WhoWhat they askWhat it works with
App developer“Can my app call this model with the client code it already uses?”The server's OpenAI-compatible endpoints
Platform team“How do we run dozens of models on one Kubernetes cluster?”Custom resources, autoscaling and canary rollouts
Performance engineer“How large should batches be before users notice the wait?”Maximum batch size and queue delay settings
5 · What it solves, and what it doesn't
solves
  • It gives a model an endpoint any app can call.
  • Batching lets one GPU answer more requests per second.
  • Autoscaling adds copies for traffic spikes and can scale to zero when idle.
  • Rollout tools let a team test a new model version on part of the traffic.
doesn't solve
  • It does not make the model more accurate; it runs whatever model it is given.
  • Bigger batches can make each individual answer slower.
  • An API key option alone is not enough to secure a vLLM server.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsOpenAI-Compatible Server, vLLM · read 28 Sept 2026
  2. repovllm-project/vllm: A high-throughput and memory-efficient inference and serving engine for LLMs, vLLM project (GitHub) · read 28 Sept 2026
  3. paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv (SOSP 2023) · read 28 Sept 2026
  4. docsBatchers, NVIDIA · read 28 Sept 2026
  5. repotriton-inference-server/server: The Triton Inference Server, NVIDIA (GitHub) · read 28 Sept 2026
  6. docsWelcome to KServe, KServe · read 28 Sept 2026
  7. docsServing Models, TensorFlow · read 28 Sept 2026