Building with AI

Token costs

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

AI model APIs charge by the token, with separate prices for the text you send in and the text the model writes back.

1 · What it is

Apps often reach an AI model through an API, a service they call over the internet. When you do this, you pay for tokens. A token is a small piece of text, such as a short word or part of a long word. In Gemini, one token works out to roughly four characters. The bill counts tokens both ways: the text you send in (input) and the text the model writes back (output).

Those two directions have different prices. In one Anthropic price table, output costs five times as much as input. Reasoning models, which work through a problem step by step, add a twist. They think in hidden tokens before answering, and those are billed as output. Gemini’s price page says the same thing: its output price includes thinking tokens.

Picture a homework helper app that sends the same long set of rules with every question. That repeated start can be cached, which means stored for reuse. Saving it to the cache costs a little more than normal input: 1.25 times the base price for a five-minute cache. Reading it back costs about a tenth of the normal input price on many Claude models. So caching pays off after one reuse. Some jobs can wait a day for an answer. Those can go through a batch service, which runs them later, at half price. Tools such as web search can add their own charges on top.

2 · Why it exists

A bill for an AI app can be hard to predict.

Words are not tokensThe price is set per token, not per word, so a word count only gives a rough guess.
Two different ratesText the model writes is often priced higher than text you send it.
Hidden thinkingSome models think in tokens you never see, and you still pay for them.
3 · How it works

Follow one request from the prompt to the bill.

Each kind of token has its own rate. The bill is the sum of every count times its rate.
  1. 1 · splitA tokenizer cuts your prompt into tokens, which are small pieces of text such as short words or parts of long words.
  2. 2 · countThe API counts the input tokens you sent and the output tokens the model wrote.
  3. 3 · priceEach count is multiplied by its own price, usually quoted per million tokens.
  4. 4 · addExtras such as cached input or batch discounts change those rates before everything is added up.

The bill is count times rate, done separately for input and output.

4 · Where it's used
WhoWhat they askWhat it works with
App developer“How many tokens will this prompt use before I send it?”A free token-counting endpoint
Support chatbot team“Can we stop paying full price for the same long instructions every time?”Caching the same opening instructions
Data team“Can we run ten thousand summaries overnight for less?”A batch job with a 24-hour window
Engineering lead“Why did the reasoning model cost more for a short answer?”Reasoning tokens billed as output
5 · What it solves, and what it doesn't
solves
  • Counting tokens before you send a request lets you estimate its cost.
  • Caching a repeated prompt start makes later reads of it much cheaper.
  • Batch jobs cut the token price in half when you can wait for results.
doesn't solve
  • A count made with one model's tokenizer may not match another model.
  • Reasoning tokens are hidden, yet they still add to the output bill.
  • A request can run out of room and stop, and you still pay for the tokens used.
  • Prices change, so any number you write down needs checking against the live price page.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsUnderstand and count tokens, Google AI for Developers · read 28 Sept 2026
  2. docsPricing, Anthropic · read 28 Sept 2026
  3. docsPrompt caching, Anthropic · read 28 Sept 2026
  4. docsReasoning models, OpenAI · read 28 Sept 2026
  5. docsBatch API, OpenAI · read 28 Sept 2026
  6. docsToken counting, Anthropic · read 28 Sept 2026
  7. docsGemini Developer API pricing, Google AI for Developers · read 28 Sept 2026