What is AWS Trainium?
AWS Trainium is a chip family designed by AWS for training and inference of AI models. The product page names the generations Trainium, Trainium2, and Trainium3. The chips power the Amazon EC2 Trn instances.
With Trn3 UltraServers, which AWS made generally available on 2 December 2025, up to 144 Trainium3 chips can be combined into one system. For very large training runs AWS references EC2 UltraClusters 3.0, which scale to hundreds of thousands of chips.
Core Features
- Three generations: Trainium, Trainium2, and Trainium3 are named on the product page
- Trainium3: AWS’s first 3nm AI chip, with 2.52 petaflops of FP8 per chip, 144 GB of HBM3e, and 4.9 TB/s of memory bandwidth
- NeuronCore architecture: Eight large NeuronCore compute engines with Tensor, Vector, Scalar, and GPSIMD sub-engines plus dedicated collective communication cores
- Quantisation: Hardware-accelerated W4A8 quantisation
- Trn3 UltraServers: Up to 144 chips, 362 FP8 petaflops, up to 20.7 TB of HBM3e, and 706 TB/s of aggregate memory bandwidth
- AWS Neuron SDK: Support for PyTorch, vLLM, HuggingFace Transformers, and TorchTitan without code changes, eager mode and FSDP, and the Neuron Kernel Interface (NKI)
- Orchestration: Works with Ray, Amazon EKS, AWS Batch, and SageMaker HyperPod
Typical Use Cases
Training large language models: Trn3 UltraServers and EC2 UltraClusters 3.0 address training runs that exceed a single instance.
Inference at production scale: AWS explicitly names inference for Trainium and highlights cost per token.
Distributed training in existing environments: Trainium resources plug into existing workflows through Amazon EKS, AWS Batch, Ray, and SageMaker HyperPod.
Continuing with existing frameworks: The Neuron SDK supports PyTorch, vLLM, HuggingFace Transformers, and TorchTitan without changes to the model code.
Benefits
- Substantial gains in memory capacity and bandwidth from Trainium2 to Trainium3 per the provider’s statement
- Up to 4.4x higher performance and 4x better performance per watt for Trn3 UltraServers versus Trn2 UltraServers per AWS
- Common frameworks run through the Neuron SDK without code migration
- Low-level hardware access through the Neuron Kernel Interface for custom kernels
- Fits into existing orchestration through EKS, AWS Batch, Ray, and SageMaker HyperPod
- No separate line item; billed through the EC2 Trn instances
Integration with innFactory
As an AWS Reseller, innFactory supports you with AWS Trainium: assessing whether your training and inference workloads are a fit, porting them via the AWS Neuron SDK, selecting suitable Trn instances and UltraServers, and integrating them into existing orchestration through Amazon EKS, AWS Batch, or SageMaker HyperPod.
Typical Use Cases
Technical Specifications
Frequently Asked Questions
What is AWS Trainium?
AWS Trainium is a chip family designed by AWS for AI workloads. The product page names the generations Trainium, Trainium2, and Trainium3. The chips power the Amazon EC2 Trn instances; the product page references Trn1 and Trn1n instances among others.
What distinguishes Trainium3?
AWS describes Trainium3 as its first 3nm AI chip and as its fourth-generation AI chip. Per the announcement each chip delivers 2.52 petaflops of FP8 compute and carries 144 GB of HBM3e with 4.9 TB/s of memory bandwidth, which AWS states is 1.5x the memory capacity and 1.7x the bandwidth of Trainium2. The product page names eight large NeuronCore compute engines with four specialised sub-engines (Tensor, Vector, Scalar, GPSIMD), hardware-accelerated W4A8 quantisation, and dedicated collective communication cores.
What are Trn3 UltraServers?
Per AWS, Trn3 UltraServers scale up to 144 Trainium3 chips, reaching 362 FP8 petaflops. Fully configured, AWS names up to 20.7 TB of HBM3e and 706 TB/s of aggregate memory bandwidth. Compared with Trn2 UltraServers, AWS states up to 4.4x higher performance, 3.9x higher memory bandwidth, and 4x better performance per watt. AWS made Trn3 UltraServers generally available on 2 December 2025.
How are Trainium chips programmed?
Through the AWS Neuron SDK. Per the product page it supports frameworks such as PyTorch, vLLM, HuggingFace Transformers, and TorchTitan without modification. AWS also names eager mode and FSDP for researchers and the Neuron Kernel Interface (NKI) for low-level hardware access.
Which AWS services does Trainium work with?
The product page names Ray, Amazon EKS, AWS Batch, and SageMaker HyperPod. For very large training runs AWS references EC2 UltraClusters 3.0, which scale to hundreds of thousands of chips.
How is Trainium billed?
Trainium has no separate line item. Billing runs through the prices of the respective Amazon EC2 Trn instances.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of AWS (official documentation). This page does not represent an offer by AWS.