AWS Inferentia
AWS Inferentia is a family of computer chips that Amazon designed to run trained AI models cheaply and quickly in its cloud.
AWS Inferentia is a chip that Amazon Web Services (AWS) designed for one job: inference. Inference means using a model that has already been trained, for example to answer a chatbot question. AWS offers it through Amazon EC2 cloud computers, called instances, that have the chips inside.
There are two generations so far. The first chip powers Inf1 instances, which hold up to sixteen chips. Inferentia2 powers Inf2 instances, which hold up to twelve. Each Inferentia2 chip has two NeuronCores, its small processing units, and 32 GiB of fast memory for the model.
The chips need their own software, a kit called AWS Neuron. Its compiler, a program that translates code, turns your model into a file the NeuronCores understand. AWS reports Inf2 reaching up to four times the throughput of Inf1, but that is a best-case vendor figure. Test your own model before you switch.
Running a trained model for real users brings three problems.
Follow one trained model onto an Inferentia chip.
- 1 · trainYou start with a model already trained in a framework such as PyTorch.
- 2 · compileThe Neuron compiler turns the model into a NEFF file, a format the chip's NeuronCores can run.
- 3 · loadThe Neuron Runtime loads that file and runs it on the NeuronCores.
- 4 · serveYour app sends requests to an EC2 Inf2 instance, which holds up to 12 Inferentia2 chips.
Inferentia does not run your code as it is: the Neuron compiler must translate the model first.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot team | “Can we serve our language model with vLLM on Inferentia?” | A large language model served through vLLM Neuron |
| Online shop | “Can product recommendations come back fast for every visitor?” | A recommendation model on Inf1 instances |
| Media app | “Can we generate images for users at lower cost?” | An image generation model on Inf2 instances |
- It gives AWS customers chips built only for inference, the step where a trained model answers.
- Inf2 instances can spread a very large model across several chips.
- The Neuron kit works with PyTorch and JAX, so teams keep much of their existing code.
- Neuron can store a model's numbers in smaller formats automatically.
- For training, AWS points to its separate Trainium chips.
- AWS offers it through Amazon EC2 instances in its cloud.
- Models must go through the Neuron compiler, so code written for other chips does not run unchanged.
- AWS's headline figures compare Inf2 with Inf1 and other EC2 instances, not promises for your model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialAWS Inferentia, Amazon Web Services · read 28 Sept 2026
- officialAmazon EC2 Inf2 Instances, Amazon Web Services · read 28 Sept 2026
- officialAmazon EC2 Inf1 Instances, Amazon Web Services · read 28 Sept 2026
- docsInferentia2 Architecture, AWS Neuron Documentation · read 28 Sept 2026
- docsInferentia Architecture, AWS Neuron Documentation · read 28 Sept 2026
- docsAWS Neuron Documentation, AWS Neuron Documentation · read 28 Sept 2026