What is AWS Inferentia?
AWS Inferentia is a chip family designed by AWS for deep learning inference. The first generation powers the Amazon EC2 Inf1 instances; the second generation, Inferentia2, powers the Inf2 instances.
The chips target cost-effective inference in production. For Inf1 instances AWS states up to 2.3x higher throughput and up to 70 percent lower cost per inference than comparable GPU instances.
Core Features
- Inferentia (first generation): Four NeuronCores and 8 GB of DDR4 per chip, support for FP16, BF16, and INT8; Inf1 instances with up to 16 chips
- Inferentia2: Two NeuronCores, 32 GB of HBM per chip, and 190 TFLOPS of FP16; Inf2 instances with up to twelve chips
- Extended data types: Inferentia2 supports FP32, TF32, and configurable FP8 among other lower-precision formats
- Distributed inference: Per AWS, Inf2 instances enable scale-out inference over ultra-high-speed connectivity between chips
- AWS Neuron SDK: Native integration with PyTorch and TensorFlow, automatic casting of high-precision models to lower-precision formats
- Framework support: PyTorch, TensorFlow, and MXNet
- Efficiency: Per AWS, up to 50 percent better performance per watt for Inferentia2 versus comparable EC2 instances
Typical Use Cases
Large language models: AWS explicitly describes Inferentia2 as optimised for large language models and latent diffusion models.
Language and text processing: Natural language processing, language translation, text summarisation, and speech recognition.
Image and video generation: Generative image and video workloads in inference.
Personalisation and fraud detection: Recommendation systems and fraud detection with high request volumes.
Benefits
- Up to 2.3x higher throughput and up to 70 percent lower cost per inference for Inf1 versus comparable GPU instances per the provider’s statement
- Substantial gains in throughput and latency from Inferentia to Inferentia2 per AWS
- Four times the memory capacity per chip for Inferentia2 compared with Inferentia
- Up to 50 percent better performance per watt for Inferentia2 per the provider’s statement
- Common frameworks run through the Neuron SDK without changing ecosystems
- No separate line item; billed through the EC2 Inf instances
Integration with innFactory
As an AWS Reseller, innFactory supports you with AWS Inferentia: assessing your inference workloads, choosing between Inf1 and Inf2 instances, porting models via the AWS Neuron SDK, and measuring throughput, latency, and cost against your current environment.
Typical Use Cases
Technical Specifications
Frequently Asked Questions
What is AWS Inferentia?
AWS Inferentia is a chip family designed by AWS for deep learning inference. The first generation powers the Amazon EC2 Inf1 instances, and Inferentia2 powers the Inf2 instances.
What specifications does AWS state for the first generation?
Per the product page, a first-generation Inferentia chip has four NeuronCores and 8 GB of DDR4 memory and supports the FP16, BF16, and INT8 data types. Inf1 instances offer up to 16 Inferentia chips. AWS states up to 2.3x higher throughput and up to 70 percent lower cost per inference than comparable GPU instances.
What distinguishes Inferentia2?
Per the product page, Inferentia2 delivers up to 4x higher throughput and up to 10x lower latency compared with Inferentia. Each chip has two NeuronCores, 32 GB of HBM memory, which is four times that of Inferentia, and 190 TFLOPS of FP16 performance. Supported data types include FP32, TF32, and configurable FP8 plus further lower-precision formats. AWS states up to 50 percent better performance per watt versus comparable EC2 instances.
How many chips do Inf2 instances offer?
Per the product page, Inf2 instances offer up to twelve Inferentia2 chips and enable scale-out distributed inference with ultra-high-speed connectivity between the chips.
How are Inferentia chips programmed?
Through the AWS Neuron SDK. Per the product page it integrates natively with popular ML frameworks such as PyTorch and TensorFlow, automatically casts high-precision models to lower-precision formats, and optimises performance. The page names PyTorch, TensorFlow, and MXNet as supported frameworks.
Which workloads suit Inferentia?
The product page names natural language processing, language translation, text summarisation, video and image generation, speech recognition, personalisation, and fraud detection. For Inferentia2, AWS highlights large language models and latent diffusion models.
Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of AWS (official documentation). This page does not represent an offer by AWS.