Products

AWS Inferentia

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

AWS Inferentia is a family of computer chips that Amazon designed to run trained AI models cheaply and quickly in its cloud.

1 · What it is

AWS Inferentia is a chip that Amazon Web Services (AWS) designed for one job: inference. Inference means using a model that has already been trained, for example to answer a chatbot question. AWS offers it through Amazon EC2 cloud computers, called instances, that have the chips inside.

There are two generations so far. The first chip powers Inf1 instances, which hold up to sixteen chips. Inferentia2 powers Inf2 instances, which hold up to twelve. Each Inferentia2 chip has two NeuronCores, its small processing units, and 32 GiB of fast memory for the model.

The chips need their own software, a kit called AWS Neuron. Its compiler, a program that translates code, turns your model into a file the NeuronCores understand. AWS reports Inf2 reaching up to four times the throughput of Inf1, but that is a best-case vendor figure. Test your own model before you switch.

2 · Why it exists

Running a trained model for real users brings three problems.

Answering costs moneyAWS says running models, called inference, is often most of the infrastructure bill for machine learning apps.
Models keep growingThe models behind AI apps are getting more complex, which pushes compute costs up.
Big models need roomLarge language models can have hundreds of billions of parameters, the numbers a model learns. Inf2 can spread such a model across several chips.
3 · How it works

Follow one trained model onto an Inferentia chip.

How a trained model runs on AWS Inferentia A trained PyTorch or JAX model goes into the Neuron compiler, the key step, which produces a NEFF file. The Neuron Runtime loads the file onto NeuronCores in the Inferentia2 chips of an EC2 Inf2 instance. User requests go in and answers come back. FROM TRAINED MODEL TO ANSWERS ON INFERENTIA Trained model Neuron compiler NEFF file Neuron Runtime Inf2 instance Your app and its users PyTorch or JAX model graph translates the model for NeuronCores ready-to-run file for NeuronCores loads the file, manages memory up to 12 chips, 2 cores each questions in, answers out requests answers
  1. 1 · trainYou start with a model already trained in a framework such as PyTorch.
  2. 2 · compileThe Neuron compiler turns the model into a NEFF file, a format the chip's NeuronCores can run.
  3. 3 · loadThe Neuron Runtime loads that file and runs it on the NeuronCores.
  4. 4 · serveYour app sends requests to an EC2 Inf2 instance, which holds up to 12 Inferentia2 chips.

Inferentia does not run your code as it is: the Neuron compiler must translate the model first.

4 · Where it's used
WhoWhat they askWhat it works with
Chatbot team“Can we serve our language model with vLLM on Inferentia?”A large language model served through vLLM Neuron
Online shop“Can product recommendations come back fast for every visitor?”A recommendation model on Inf1 instances
Media app“Can we generate images for users at lower cost?”An image generation model on Inf2 instances
5 · What it solves, and what it doesn't
solves
  • It gives AWS customers chips built only for inference, the step where a trained model answers.
  • Inf2 instances can spread a very large model across several chips.
  • The Neuron kit works with PyTorch and JAX, so teams keep much of their existing code.
  • Neuron can store a model's numbers in smaller formats automatically.
doesn't solve
  • For training, AWS points to its separate Trainium chips.
  • AWS offers it through Amazon EC2 instances in its cloud.
  • Models must go through the Neuron compiler, so code written for other chips does not run unchanged.
  • AWS's headline figures compare Inf2 with Inf1 and other EC2 instances, not promises for your model.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialAWS Inferentia, Amazon Web Services · read 28 Sept 2026
  2. officialAmazon EC2 Inf2 Instances, Amazon Web Services · read 28 Sept 2026
  3. officialAmazon EC2 Inf1 Instances, Amazon Web Services · read 28 Sept 2026
  4. docsInferentia2 Architecture, AWS Neuron Documentation · read 28 Sept 2026
  5. docsInferentia Architecture, AWS Neuron Documentation · read 28 Sept 2026
  6. docsAWS Neuron Documentation, AWS Neuron Documentation · read 28 Sept 2026