Token costs
AI model APIs charge by the token, with separate prices for the text you send in and the text the model writes back.
Apps often reach an AI model through an API, a service they call over the internet. When you do this, you pay for tokens. A token is a small piece of text, such as a short word or part of a long word. In Gemini, one token works out to roughly four characters. The bill counts tokens both ways: the text you send in (input) and the text the model writes back (output).
Those two directions have different prices. In one Anthropic price table, output costs five times as much as input. Reasoning models, which work through a problem step by step, add a twist. They think in hidden tokens before answering, and those are billed as output. Gemini’s price page says the same thing: its output price includes thinking tokens.
Picture a homework helper app that sends the same long set of rules with every question. That repeated start can be cached, which means stored for reuse. Saving it to the cache costs a little more than normal input: 1.25 times the base price for a five-minute cache. Reading it back costs about a tenth of the normal input price on many Claude models. So caching pays off after one reuse. Some jobs can wait a day for an answer. Those can go through a batch service, which runs them later, at half price. Tools such as web search can add their own charges on top.
A bill for an AI app can be hard to predict.
Follow one request from the prompt to the bill.
- 1 · splitA tokenizer cuts your prompt into tokens, which are small pieces of text such as short words or parts of long words.
- 2 · countThe API counts the input tokens you sent and the output tokens the model wrote.
- 3 · priceEach count is multiplied by its own price, usually quoted per million tokens.
- 4 · addExtras such as cached input or batch discounts change those rates before everything is added up.
The bill is count times rate, done separately for input and output.
| Who | What they ask | What it works with |
|---|---|---|
| App developer | “How many tokens will this prompt use before I send it?” | A free token-counting endpoint |
| Support chatbot team | “Can we stop paying full price for the same long instructions every time?” | Caching the same opening instructions |
| Data team | “Can we run ten thousand summaries overnight for less?” | A batch job with a 24-hour window |
| Engineering lead | “Why did the reasoning model cost more for a short answer?” | Reasoning tokens billed as output |
- Counting tokens before you send a request lets you estimate its cost.
- Caching a repeated prompt start makes later reads of it much cheaper.
- Batch jobs cut the token price in half when you can wait for results.
- A count made with one model's tokenizer may not match another model.
- Reasoning tokens are hidden, yet they still add to the output bill.
- A request can run out of room and stop, and you still pay for the tokens used.
- Prices change, so any number you write down needs checking against the live price page.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsUnderstand and count tokens, Google AI for Developers · read 28 Sept 2026
- docsPricing, Anthropic · read 28 Sept 2026
- docsPrompt caching, Anthropic · read 28 Sept 2026
- docsReasoning models, OpenAI · read 28 Sept 2026
- docsBatch API, OpenAI · read 28 Sept 2026
- docsToken counting, Anthropic · read 28 Sept 2026
- docsGemini Developer API pricing, Google AI for Developers · read 28 Sept 2026