NPUsConcepts

Neural processing units

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An NPU is a part of a chip built to speed up AI models at low power.

1 · What it is

A neural processing unit, or NPU, is a part of a chip made for one job: running the maths inside AI models. Your phone or laptop also has a CPU, the main all-purpose processor, and a GPU, the graphics chip. The NPU is the specialist next to them. Qualcomm says its NPU is a little harder to program, but in return it is built for top performance and low power.

Why a special chip? The operations inside AI are the multiplications, additions and other small steps that AI needs in huge numbers. Microsoft says the NPU works on lots of data in parallel, doing trillions of these steps per second. Microsoft adds that for AI jobs it wastes less energy than a CPU or GPU would, so the battery lasts longer.

Now follow one example from start to finish. Picture an app that writes captions for what people say on a video call. The app is made up, but it follows the steps Microsoft describes for Windows. The captions run all call long, so battery life matters.

Step 1 is to convert the model so it fits the NPU. A model is trained (taught from examples), and models are usually trained with numbers in a large format called FP32. Many NPUs can only do maths on smaller whole numbers, such as INT8, because that is faster and saves power. So the model has to be converted, which is called quantizing. Many models are already available in a converted, ready-to-use form, so this step is often done for you.

Step 2 is to pick the right chip. Microsoft says the recommended way to run AI on the NPU in Windows is a tool called Windows ML. Windows ML looks for the chips the computer has and picks the fastest option it finds. If that option fails or is missing, it falls back to the GPU or CPU.

Step 3 is the NPU doing the maths. Intel says its NPU has special blocks for AI steps such as multiplying big grids of numbers. Intel’s NPU keeps its work in small, fast memory on the chip itself, and moves data to the computer’s main memory as little as possible, so it does more work for each bit of power.

Step 4 is the caption text going back to the app. The work happened on the laptop, not on a server. Keeping the work on the device also helps privacy, because the data stays there.

Where will you find NPUs? Apple’s first Neural Engine shipped in the A11 chip of the iPhone X in 2017. Face ID (unlocking the phone with your face) and Memoji (animated emoji that copy your face) ran on it, right on the phone. Intel puts an NPU inside its Core Ultra processors. Copilot+ PCs are a newer kind of Windows 11 computer, each built around a fast NPU. That NPU is rated above 40 trillion operations per second, written as TOPS.

What does a TOPS number mean for you? So a rating of “40 TOPS” means the chip can do over 40 trillion of those steps each second. Many of the newest Windows AI tools only work with an NPU that fast.

The NPU does not replace the CPU and GPU; the three chips share the work. When the Neural Engine takes the heaviest AI work, the CPU and GPU are left free for other jobs. An app that ignores the NPU gets none of that benefit.

2 · Why it exists

Running AI on an ordinary processor is possible, but it costs speed and battery.

AI is heavy mathsMachine learning needs a huge number of multiplications and additions. An NPU is less flexible, but it runs these sums with less power.
Batteries run downFor AI features that run for a long time, where battery life matters, an NPU is the best option.
Busy main chipsIf the main processor and graphics chip do all the AI work, they have less room for everything else.
3 · How it works

Follow a made-up captioning app, built with Windows ML, as it runs on a laptop.

The highlighted box is step 3: the NPU runs the model's maths. The CPU and GPU remain free for other work.
  1. 1 · convertThe model is converted (quantized) to fit the NPU: its numbers change from a big format, FP32, to small whole numbers, INT8. Many models come already converted.
  2. 2 · routeWindows ML checks which chips the laptop has and picks the fastest option it finds. If the NPU is missing, it uses the GPU or CPU instead.
  3. 3 · computeThe NPU's dedicated blocks work side by side on common AI maths, like multiplying big grids of numbers (matrix multiplication) and sliding filters over images (convolution).
  4. 4 · returnThe result goes back to the app. The work was done on the device, not on a server.

An NPU is about efficiency: it runs AI tasks using less power than a CPU or GPU, and frees them for other work.

4 · Where it's used
WhoWhat they askWhat it works with
Phone owner“Which chip on my iPhone runs Face ID?”The Neural Engine, which powers on-device features such as Face ID
Laptop buyer“Does this laptop's NPU meet the level needed for Copilot+ PC features?”The NPU's rated trillions of operations per second
App developer“Will my model run on the NPU, or does it need converting first?”The model's number format and the device's supported formats
Video call user“Which chip should run a feature that stays on for a whole call?”The NPU, the best option when battery life matters for sustained use
5 · What it solves, and what it doesn't
solves
  • NPUs run AI tasks with less power than a CPU or GPU.
  • They free the CPU and GPU for other tasks.
  • They let AI features run on the device, so the data can stay there.
  • They suit AI features that run for a long time, where battery life matters.
doesn't solve
  • An NPU cannot use a model until software is written or converted to target it.
  • Many NPUs support only low-precision integer maths, so models often need converting first.
  • An NPU is not a replacement for the CPU and GPU; the three work together.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsCopilot+ PCs developer guide, Microsoft Learn · read 28 Sept 2026
  2. officialAll about neural processing units (NPUs), Microsoft Support · read 28 Sept 2026
  3. officialDeploying Transformers on the Apple Neural Engine, Apple Machine Learning Research · read 28 Sept 2026
  4. officialUnlocking on-device generative AI with an NPU and heterogeneous computing, Qualcomm Technologies · read 28 Sept 2026
  5. docsQuick overview of Intel's Neural Processing Unit (NPU), Intel · read 28 Sept 2026