Concepts

AI accelerators

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

AI accelerators are chips built to speed up the maths inside AI models, which is largely multiply-add sums.

1 · What it is

An AI accelerator is a chip designed to speed up the maths of AI models. Four families appear here: graphics cards (GPUs), Google’s tensor processing units (TPUs), neural processing units (NPUs) and Amazon’s Trainium and Inferentia.

The one sum AI keeps doing

Picture a photo app deciding that a picture shows a dog. The photo becomes a grid of numbers, and the model works through layers. A neural network is a model built from these layers of simple number steps. In each layer it takes a list of numbers, multiplies each one by a number it learned in training, and adds the results into a total. NVIDIA calls this multiply-add the most common operation in modern neural networks.

Here is a made-up example with tiny numbers. Say three inputs are 2, 5 and 1, and the learned numbers are 3, 1 and 4. The layer works out 2 times 3, plus 5 times 1, plus 1 times 4. That is 6 plus 5 plus 4, so the total is 15.

Many rows of these form a grid called a matrix. At the end, “dog” gets the highest score.

An accelerator is hardware built to run many of these sums at the same time.

Learning versus answering

AWS says it built Trainium to make fast AI training and inference cheaper at large scale. AWS made Inferentia chips to run trained AI models. AWS says Inferentia aims to be fast while keeping cost as low as it can on its cloud computers.

GPUs and TPUs

A GPU is built to do lots of work in parallel, which means side by side. It stays busy by running many threads at once. A thread is one small stream of work. NVIDIA added Tensor Cores to its GPUs. These are special parts of the chip that speed up grid-shaped multiply-and-add work.

Google’s TPUs are custom chips built to speed up machine learning, the way computers learn from examples. Their matrix unit, the part that does grid sums, is a 128 by 128 grid of multiply-add hardware. The software cuts big matrix jobs into blocks that fit it.

Think of a giant times-table sheet that is too big for one desk. You cut it into square pieces, hand each piece to a helper, and then combine their answers. Google describes the TPU software doing something like this: it tiles a big matrix multiply into smaller blocks so the matrix unit can work through them efficiently.

Google’s first TPU, described in a 2017 paper, was made to speed up inference. Compared with a CPU and GPU of the time, on Google’s own neural network work, the TPU was on average about 15 to 30 times faster. It also did far more work for each unit of electricity used.

NPUs in PCs

Accelerators also sit inside personal computers. Microsoft’s Copilot+ PCs are built around an NPU. The NPU uses less energy for AI work than a CPU or GPU, so batteries last longer.

Limits

An NPU needs software written to use it. Some models fit a chip badly. A batch is a group of inputs handled together; its shape is the size of its number grids. A TPU model is prepared, or compiled, for batches of one shape. If later batches come in a different shape, it stops working.

Each step has two parts: fetching numbers from memory and doing the sums. If fetching takes longer than the sums, the job is memory limited. Then the sum hardware spends its time waiting for numbers to arrive.

2 · Why it exists

AI models ask computers to repeat one kind of sum an enormous number of times.

One sum, repeatedNVIDIA's guide says multiply-add, multiplying two numbers and adding the result to a total, is the most common operation in modern neural networks.
Sums come in gridsGoogle says TPUs use hardware designed for the large matrix operations common in machine learning.
Power and batteryMicrosoft says an NPU uses less energy for AI work than a CPU or GPU would, so a laptop battery lasts longer.
3 · How it works

Follow one photo through one layer of an image model.

Different accelerators, one shared trick: dedicated hardware for blocks of multiply-add sums.
  1. 1 · numbersThe input, such as a photo, is turned into numbers.
  2. 2 · matrixA model layer can be seen as many rows of multiply-adds, each row matching a list of inputs against a list of learned numbers.
  3. 3 · blocksThe software splits a big matrix multiply into smaller blocks that fit the chip's matrix unit.
  4. 4 · parallelThe chip runs many of those block calculations at once, as many small parallel jobs.
  5. 5 · outputThe results pass to the next layer until the model produces its answer.

The speed comes from doing many simple sums at once, not from doing one sum cleverly.

4 · Where it's used
WhoWhat they askWhat it works with
Researcher training a model“Which chips should I rent to train this model faster?”Cloud AI chips such as TPUs or AWS Trainium
App company serving a chatbot“How do we answer millions of questions without the bill exploding?”Inference chips such as AWS Inferentia
PC maker“Can our AI features run on the PC without draining the battery?”An NPU, as in Copilot+ PCs
Student with a gaming PC“Can I run a small image model at home?”A GPU or a CPU, depending on the job
5 · What it solves, and what it doesn't
solves
  • It speeds up the repeated multiply-add sums in AI models.
  • Google's first TPU did far more work for the same electricity than the CPUs and GPUs of its day.
  • Chips built for answering, such as Inferentia, aim to serve AI answers at lower cost.
  • An NPU lets a PC run AI features on the device itself.
doesn't solve
  • An NPU needs software written to use it.
  • Some models fit badly, for example ones whose data sizes keep changing on a TPU.
  • An NPU works alongside the CPU and GPU rather than replacing them.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsIntroduction to Cloud TPU, Google Cloud · read 28 Sept 2026
  2. docsGPU Performance Background User's Guide, NVIDIA · read 28 Sept 2026
  3. docsDevelop AI applications for Copilot+ PCs, Microsoft · read 28 Sept 2026
  4. officialAWS Trainium, Amazon Web Services · read 28 Sept 2026
  5. officialAWS Inferentia, Amazon Web Services · read 28 Sept 2026
  6. paperIn-Datacenter Performance Analysis of a Tensor Processing Unit, arXiv · read 28 Sept 2026