Batch inference
Batch inference sends many AI requests as one job that runs in the background, trading an instant reply for a lower price.
Normally an app sends one request to an AI model and waits for the reply. That makes sense for a chat. It makes less sense when you have ten thousand product descriptions to write and nobody is watching the screen.
Batch inference is the other way to ask. You gather your requests into a single job. The provider works through it in the background. Later you download a file of answers. Because you are not waiting, OpenAI, Anthropic and Google charge half the usual price for this kind of work.
The catch is time. OpenAI says a batch is done within 24 hours, and Anthropic says most of its batches finish in under an hour. At Anthropic, a batch that has not finished after 24 hours expires. So batch inference fits jobs where nobody is waiting: testing a model on many questions, sorting large piles of posts, or embedding a library of documents. It is not for a live conversation.
For engineers: at OpenAI, requests go in as a JSON Lines file, one request per line, each tagged with a unique custom_id. The output file may come back in a different order, so you match answers to requests by that id, not by position. An Anthropic batch can hold up to 100,000 requests or 256 MB. Its results can be downloaded for 29 days. On Amazon Bedrock you collect the output files from S3 storage. On Google Cloud, cache discounts and the batch discount do not stack.
Sending thousands of requests one at a time is slow and costly.
Follow one file of requests from upload to answers.
- 1 · prepareAt OpenAI, you write your requests into one .jsonl file, one request per line, each with a unique custom_id.
- 2 · submitYou start a batch job, then check its status.
- 3 · processThe provider works through the requests in the background, each handled on its own.
- 4 · collectWhen the job ends, you download the results and match each answer to its request by id.
Batch inference is for work that can wait.
| Who | What they ask | What it works with |
|---|---|---|
| Evaluation team | “How does the new prompt score on our thousands of test cases?” | A file of test prompts run as one batch |
| Trust and safety | “Which of yesterday's posts break our rules?” | A large set of user posts to classify |
| Search team | “Can we embed our whole document library?” | A document library to embed |
| Online shop | “Can we draft descriptions for all our products?” | The product catalogue, one request per item |
- Batched requests cost half the standard price at OpenAI, Anthropic and the Gemini API.
- At OpenAI, batches have their own rate limits, separate from normal requests.
- One job replaces a long queue of separate calls.
- It is not for chat or anything a person is waiting on.
- At Anthropic, a batch expires if it has not finished within 24 hours.
- On Amazon Bedrock, batch jobs do not support tool calling or structured output.
- Each request runs alone, so there is no back-and-forth with the model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsBatch API, OpenAI · read 28 Sept 2026
- docsBatch processing, Anthropic · read 28 Sept 2026
- docsBatch API, Google AI for Developers · read 28 Sept 2026
- docsBatch inference with Gemini, Google Cloud · read 28 Sept 2026
- docsProcess multiple prompts with batch inference, Amazon Web Services · read 28 Sept 2026