Tensor Processing Units
TPUs are chips Google designed to speed up the maths of machine learning, rented out through Google Cloud.
A TPU, short for Tensor Processing Unit, is a computer chip Google built for the maths inside neural networks. Google says a TPU cannot run a word processor. What it does well is multiply huge grids of numbers, called matrices, very fast.
Why so much multiplying?
Picture a tiny network layer with three inputs and two neurons (small maths units). Each neuron multiplies every input by a weight, a number that says how much that input matters, then adds the results. Say the inputs are 2, 1 and 3. With weights 1, 0 and 2, the first neuron gets 2×1 + 1×0 + 3×2 = 8. With weights 0, 1 and 1, the second gets 4. Six multiplications, then two sums. Real models do this at huge scale, and Google says it is often the biggest job when a finished model runs.
The memory problem
A CPU, the all-rounder chip, fetches numbers from memory, works on them, then writes the answer back, every time. Memory is slow next to the maths. A GPU has thousands of small maths units working at once, but it still reads and stores partial results in small on-chip storage for every sum.
How a TPU gets around it
Think of a bucket brigade: nobody runs back to the well; buckets pass hand to hand. A TPU’s matrix unit is a grid of multiply-accumulators, tiny units that multiply two numbers and add the result to a running total. Each passes its total straight to its neighbour, so no memory trip is needed during the multiplication.
A translator program, the XLA compiler, turns the model’s maths into chip instructions. The host computer feeds data in, the TPU stores it in its own fast memory, runs it through the grid, and sends results back.
Inside the chip
Each TensorCore, the chip’s main work area, has three parts. The matrix unit (MXU) does the big multiplications. The vector unit handles steps like activations, which decide how strongly each neuron “fires”. The scalar unit handles control flow, meaning what to do next. Multiplications take inputs in a short number format called bfloat16, while running totals use a longer format, FP32.
The grid has a fixed size, so the compiler cuts big jobs into smaller blocks called tiles. If a job can’t fill the grid, it pads the gaps with zeros, which wastes part of the chip.
Where TPUs came from
In 2013, Google saw that neural networks might mean it needed twice as many data centers. It has used TPUs in its data centers since 2015 to run trained networks. In a Google study, the first TPU ran already-trained networks about 15 to 30 times faster, on average, than a GPU or CPU of the same era. Today TPUs power Gemini, Search, Photos and Maps, and they can be rented through Google Cloud. Code written in JAX or PyTorch, two popular AI toolkits, can run on them.
When to pick which
Google points to TPUs for models that are mostly matrix maths, large models trained for weeks, and models with huge embeddings (tables that turn each item into a list of numbers). It points to CPUs for quick experiments and small models. TPUs are a poor fit when the main training loop has its own custom operations.
Limits
Programs full of “if this, do that” branching are a poor match, as is work needing very precise arithmetic. The compiler prepares the model for the first batch of data, so models whose data shapes keep changing are a poor fit.
Ordinary processors are not built around the maths of neural networks.
Follow one layer's multiplication through a TPU.
- 1 · compileA translator program called the XLA compiler turns the model's maths into machine code, the low-level instructions the TPU chip actually runs.
- 2 · streamThe host computer sends data in a steady flow to an infeed queue (a waiting line of data), and the TPU stores it in high-bandwidth memory, its own very fast memory.
- 3 · multiplyInside the matrix unit, each multiply result passes straight to the next multiply-accumulator (a tiny unit that multiplies two numbers and adds the result to a running total), with no memory access along the way.
- 4 · returnFinished results go into an outfeed queue, a waiting line the host computer reads back from.
The speed comes from not visiting memory during the multiplication.
| Who | What they ask | What it works with |
|---|---|---|
| Research lab | “Can we train a very large model for several weeks?” | Large models with large batches that train for weeks |
| Recommendation team | “Can our ranking model handle huge lookup tables of item features?” | Models with ultra-large embeddings (huge tables that turn each item into a list of numbers) |
| Student with a small model | “Do I need a TPU for my quick experiment?” | Usually a CPU, which suits quick prototyping and small models |
| App developer | “Can I run my PyTorch or JAX model on Google's chips?” | Cloud TPUs, which support PyTorch and JAX |
- Large matrix multiplications run fast because partial results stay inside the chip.
- On-chip high-bandwidth memory lets you use larger models and batch sizes.
- Many chips can be joined by fast links into slices and pods (groups of chips wired together) for very large jobs.
- Common frameworks such as JAX and PyTorch can target TPUs.
- A TPU cannot run everyday software like a word processor.
- Programs with lots of branching or many element-wise operations are a poor fit.
- Workloads that need high-precision arithmetic are not suited to TPUs.
- Models whose tensor shapes keep changing are not well suited to TPUs.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsIntroduction to Cloud TPU, Google Cloud · read 28 Sept 2026
- docsTPU architecture, Google Cloud · read 28 Sept 2026
- officialTensor Processing Units (TPUs), Google Cloud · read 28 Sept 2026
- paperIn-Datacenter Performance Analysis of a Tensor Processing Unit, arXiv · read 28 Sept 2026
- officialAn in-depth look at Google's first Tensor Processing Unit (TPU), Google Cloud · read 28 Sept 2026