CPUsConcepts

Central processing units

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A CPU is the chip that runs a computer's software, carrying out instructions with just a few cores backed by lots of cache memory.

1 · What it is

A CPU, short for central processing unit, is the chip that runs a computer’s software. People often call it simply the processor. It carries out instructions, does sums and moves data around.

Cores and threads

A CPU has just a few cores backed by lots of cache memory. Cache is a small, fast store of memory close to the cores. A core is one physical worker inside the chip. A thread is one stream of instructions, like one to-do list, and one core may juggle a single thread or several. A CPU is built to race through steps one after another, and runs only a few dozen threads at once.

Today most CPUs are multi-core: several cores share one chip, so different jobs can run at the same time.

The instruction cycle

Every instruction goes through a repeating loop. Fetch pulls the next instruction out of memory. Decode works out what it means and what it needs. Execute does the operation. Store writes the result to memory or a register, an ultra-fast spot for temporary data.

The control unit directs the other parts. The arithmetic logic unit, or ALU, does sums and logic checks. A clock keeps every step in time. The cache, in levels called L1, L2 and L3, is fast memory near each core that holds often-used instructions and data, so the CPU avoids waiting on slower main memory, a bit like keeping your pencil case on your desk instead of in your locker.

Designers add tricks to cut waiting, such as pipelining, which overlaps the stages of the loop. A higher clock speed alone does not make a faster chip; design, cores, cache and the job matter too.

Where CPUs shine

CPUs are fast and versatile. They suit tasks where you click and type and expect a quick reply, like loading the file you asked for.

Why AI maths is different

AI depends on matrix maths: huge grids of multiplications and additions. A GPU, or graphics processing unit, is a specialist chip that helps the CPU with heavily parallel jobs like graphics and AI. Its hundreds of cores juggle thousands of threads together. Each thread is slower, but far more work gets done.

Picture a grid of sums as a big multiplication table that a model needs to fill in. Each box is one small multiply, and no box depends on its neighbour. A CPU splits the table among its few cores. Even with each core juggling more than one thread, it can run only a few dozen threads at once, so the table gets done a slice at a time, and every other slice waits in line.

A GPU takes a different path. It breaks the job into thousands of separate tasks and works on them at the same moment. Each task is a little slower, but because so many run side by side, the whole table fills in far faster. For about the same price and power, a GPU gets through much more of this kind of work.

The CPU as the organiser

That does not make the CPU useless for AI. The GPU works alongside the CPU; it does not replace it. Arm says the CPU coordinates inference, the work of using a trained model to get answers, and manages the data pipelines, the steps that move data in and out. Like a head chef, it plans the meal while helpers chop.

Sometimes the split changes. When a model is too big for the GPU’s own memory, llama.cpp can put part of it on the GPU and run the rest on the CPU.

CPUs still run AI

The open-source project llama.cpp runs language models on many chips, including Apple silicon and x86 chips with AVX instructions. It can shrink a model with integer quantization, storing numbers with fewer bits, to cut memory use. Google’s LiteRT can use the CPU, the GPU or an NPU, a chip made for AI maths. That lets open models like Gemma run on your own device. ONNX Runtime, a tool for running AI models, can use a GPU when possible and fall back to the CPU.

2 · Why it exists

AI maths is the kind of work a CPU was not built for.

Few workersA CPU has only a few cores, so it can handle only a few streams of instructions at the same moment.
Built for orderIt is built to rush through one chain of steps, not thousands of sums side by side.
Matrix-heavy AIAI leans on huge grids of multiplications, which suit GPUs better.
3 · How it works

One AI calculation on a CPU and a GPU.

The same sums on four strong cores or many simple ones.
  1. 1 · fetchThe CPU reads the next instruction and its data, keeping often-used data in a fast cache.
  2. 2 · executeEach core works through its own list of instructions, called a thread.
  3. 3 · queueWith few cores, a big grid of sums waits in line, a slice at a time.
  4. 4 · compareA GPU runs thousands of threads in parallel, taking far more sums per round.

A CPU is fast at one thing after another; AI maths rewards doing many small things at once.

4 · Where it's used
WhoWhat they askWhat it works with
Laptop owner“Can I run a chatbot without a graphics card?”A small, compressed model on the CPU
App developer“What happens if the phone has no AI accelerator?”The runtime falls back to the CPU
Server engineer“Which chip runs this model, and which is backup?”A hardware list with the CPU last
Student“Why does my training script run so slowly?”Matrix maths queued on a few cores
5 · What it solves, and what it doesn't
solves
  • It runs software and the operating system, and with tools like llama.cpp it can run AI models too.
  • Its large caches cut the wait for data it uses often.
  • It moves quickly through interactive tasks.
  • It coordinates AI work and works alongside specialised chips.
doesn't solve
  • For similar price and power, a GPU gets through far more instructions.
  • A higher clock speed alone does not guarantee better performance.
  • Big models can use a lot of memory; quantization is one way llama.cpp cuts that.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialWhat is a CPU?, Arm · read 28 Sept 2026
  2. docsCUDA Programming Guide, Introduction, NVIDIA · read 28 Sept 2026
  3. officialWhat's the Difference Between a CPU and a GPU?, NVIDIA · read 28 Sept 2026
  4. repollama.cpp, ggml-org · read 28 Sept 2026
  5. docsONNX Runtime Execution Providers, ONNX Runtime · read 28 Sept 2026
  6. docsLiteRT, Google AI for Developers · read 28 Sept 2026