AWS Trainium
AWS Trainium is a computer chip Amazon designed for training and running AI models, rented through special Amazon EC2 cloud servers.
AWS Trainium is a chip that Amazon Web Services (AWS) designed for AI work. It is built for training, which means teaching a model from data, and for inference, which means running a trained model to answer requests. You cannot buy it for a desk. You rent it inside AWS on EC2 Trn instances.
AWS made it for one job: AI work. Inside each Trainium chip are maths engines called NeuronCores. A Trn2 instance links 16 Trainium2 chips with NeuronLink, AWS’s own chip-to-chip connection, and a Trn2 UltraServer joins 64 chips across four instances. AWS sells it as one joined-up package: the chips, the servers and networking around them, plus software and services.
The software that makes this usable is the AWS Neuron SDK. It lets you keep writing in PyTorch or JAX, two popular tools for building AI models. Then it compiles your model, which means it translates it into a file the NeuronCores can run. Trainium was the second generation of machine learning chip from AWS. AWS has also made Trainium2 and Trainium3.
Training big AI models is slow, costly and needs many chips working together.
Follow one training script from your laptop to the chip.
- 1 · writeYou write or reuse model code in PyTorch or JAX, mostly changing which device the code targets, from cuda (for GPUs) to neuron.
- 2 · rentYou start an Amazon EC2 Trn instance, a cloud server that holds one or more Trainium chips.
- 3 · compileThe Neuron compiler turns your model's graph into a NEFF file the chip understands.
- 4 · runThe Neuron runtime loads that file and runs it on the NeuronCores, the maths engines inside each chip.
- 5 · scaleFor bigger jobs, NeuronLink joins chips so they share the work.
Trainium is a chip made only for AI maths. The Neuron SDK is the bridge that lets familiar code run on it.
| Who | What they ask | What it works with |
|---|---|---|
| AI start-up | “Can we pretrain our language model for less money?” | Trn instances in the AWS cloud |
| Researcher | “Will my existing PyTorch training script run here?” | Native PyTorch support in the Neuron SDK |
| Serving team | “Can we run our model for users with a standard vLLM server?” | vLLM on Trainium through Neuron |
| Performance engineer | “How do I hand-write faster chip code for one layer?” | The Neuron Kernel Interface |
- It gives AWS customers a chip built for AI training and inference.
- Chips can be joined in large groups for very large training jobs.
- Common tools like PyTorch, JAX, Hugging Face and vLLM work through the Neuron SDK.
- Experts can write custom kernels, small hand-tuned pieces of chip code, with NKI (the Neuron Kernel Interface).
- You rent it inside AWS on EC2 Trainium instances, not as a chip for your own computer.
- Code still has to go through the Neuron compiler and runtime, so some changes may be needed.
- The cost savings AWS quotes are comparisons to chosen instances, not a promise for every model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialAI Accelerator - AWS Trainium, Amazon Web Services · read 28 Sept 2026
- officialAmazon EC2 Trn1 instances, Amazon Web Services · read 28 Sept 2026
- officialAmazon EC2 Trn2 instances, Amazon Web Services · read 28 Sept 2026
- officialAWS Neuron, Amazon Web Services · read 28 Sept 2026
- docsAWS Neuron documentation, Amazon Web Services · read 28 Sept 2026
- docsTrainium architecture, Amazon Web Services · read 28 Sept 2026