GPUsConcepts

Graphics processing units

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A GPU is a chip built to run many similar calculations at the same time, which suits the matrix maths inside neural networks.

1 · What it is

A GPU, short for graphics processing unit, is a chip that does lots of simple sums at the same time. According to NVIDIA, it began as a chip for drawing 3D pictures.

Picture a thousand maths quizzes to mark. One expert teacher could mark them one by one. Or a thousand students could mark one each, all at once. When every job is small and alike, the students win.

CPU and GPU

The CPU is the computer’s main chip. A core is a part of a chip that can follow instructions by itself. AMD’s guide says a CPU usually has a few strong cores, often 4 to 64. A GPU has many simpler cores, in the hundreds or thousands.

NVIDIA explains that a GPU spends more of its tiny switches (transistors) on doing maths, while a CPU spends more on keeping data close and deciding what to do next. So, for the same price and power, a GPU does far more work per second.

Inside a GPU

A program is broken into threads: tiny tasks that each handle one small piece of the data. On an NVIDIA GPU, threads are packed into thread blocks, and each block runs on one streaming multiprocessor, or SM. The A100 has 108 SMs, so a program needs lots of blocks to keep them all busy.

Why AI needs this

A matrix is a grid of numbers. NVIDIA calls multiplying matrices a basic building block of many neural network operations. (A neural network is the kind of program behind most modern AI.)

Each number in the answer comes from a dot product: multiply pairs of numbers, then add the results together. Multiply two grids that are each 1,000 by 1,000, and you need 1,000 x 1,000 x 1,000 multiply-adds. That is one billion. Each one is easy; there are just a huge number. That is the thousand-students kind of job a GPU likes.

The copying catch

The computer’s main memory and the GPU’s memory are separate, so data has to be copied over first. On the older Tesla V100, that link tops out around 16 GB per second, while the GPU’s own memory can reach 898 GB per second. So it is smart to keep in-between results on the GPU.

A real example in code

PyTorch, a popular AI toolkit, keeps a model’s inputs, outputs and learned numbers in tensors. A tensor is much like a grid of numbers in NumPy, a common Python maths library, except that it can also live on a GPU.

Here is the catch for beginners. A new tensor is made on the CPU. It will not jump to the GPU by itself; you ask with the .to method:

import torch

x = torch.rand(1000, 1000)   # made on the CPU
x = x.to("cuda")             # copied to the GPU
y = x @ x                    # matrix multiply runs on the GPU

That middle line is a real copy. PyTorch’s guide warns that copying big tensors between devices can cost a lot of time and memory. So move the data once, do the work there, and bring back only the answer.

Limits

A GPU is not magic. AMD’s guide still suggests the CPU for work full of “if this, then that” choices. The GPU does best when every thread runs the same instruction on different data.

Small jobs are a problem too. A big NVIDIA GPU can have more than 160,000 threads active at once. Adding two short lists of numbers leaves most of them with nothing to do. NVIDIA says that when a job is too small, or cannot be split up, the chip is under-used.

And the copies never go away. NVIDIA says they are costly and should be kept to a minimum.

2 · Why it exists

Neural networks need a huge amount of the same simple maths, and a CPU is built for something else.

Few fast coresA CPU has a few powerful cores tuned to finish one line of instructions quickly, not to run thousands side by side.
Matrix maths everywhereMatrix multiplication is a basic building block of many neural network layers.
Data has to moveBefore a GPU can work, data must be copied into its memory, and those copies take time.
3 · How it works

Follow one layer of a neural network through a GPU.

The GPU wins by splitting one big matrix multiplication into many small pieces and computing them all at once.
  1. 1 · copyThe program starts on the CPU and copies the input numbers and model weights into GPU memory.
  2. 2 · splitThe matrix multiplication is cut into many tiles, and each tile is handed to a group of threads.
  3. 3 · computeThousands of threads run the same multiply-and-add instructions on different numbers at the same time.
  4. 4 · returnThe answers are kept on the GPU where possible, because copies back to the CPU are costly.

A GPU is not faster at one sum. It is faster because it does many sums at once.

4 · Where it's used
WhoWhat they askWhat it works with
Student with a laptop“Why does my image model run so much faster when I switch the notebook to a GPU?”The same matrix maths, spread across thousands of threads
Model trainer“Will my model and its training data fit in GPU memory?”The GPU's memory size and bandwidth
5 · What it solves, and what it doesn't
solves
  • Runs many similar calculations in parallel, which suits matrix multiplication.
  • Gives much higher throughput and memory bandwidth than a CPU at a similar price and power.
  • Is typically used for running trained models (inference).
doesn't solve
  • It does not make branch-heavy, step-by-step logic faster; that still suits a CPU.
  • It does not remove the cost of copying data between CPU and GPU memory.
  • Small jobs may not keep the GPU busy.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsIntroduction (CUDA Programming Guide), NVIDIA · read 28 Sept 2026
  2. docsGPU Performance Background User's Guide, NVIDIA · read 28 Sept 2026
  3. docsMatrix Multiplication Background User's Guide, NVIDIA · read 28 Sept 2026
  4. docsHIP programming model, AMD · read 28 Sept 2026
  5. docsCUDA C++ Best Practices Guide, NVIDIA · read 28 Sept 2026
  6. docsTensors, PyTorch · read 28 Sept 2026
  7. officialWhat is AI inference?, NVIDIA · read 28 Sept 2026