Skip to main content
Cloud / AWS / Products / AWS Inferentia: Chips for Deep Learning Inference

AWS Inferentia: Chips for Deep Learning Inference

AWS Inferentia is a chip family designed by AWS for deep learning inference, powering the Amazon EC2 Inf1 and Inf2 instances.

Machine Learning
Pricing Model No separate line item; billed through the prices of the Amazon EC2 Inf instances
Availability Availability varies by instance type and Region, see official documentation
Data Sovereignty Matches the Region in which the EC2 instance runs
Reliability The SLA of the service used applies, in particular Amazon EC2 SLA

What is AWS Inferentia?

AWS Inferentia is a chip family designed by AWS for deep learning inference. The first generation powers the Amazon EC2 Inf1 instances; the second generation, Inferentia2, powers the Inf2 instances.

The chips target cost-effective inference in production. For Inf1 instances AWS states up to 2.3x higher throughput and up to 70 percent lower cost per inference than comparable GPU instances.

Core Features

  • Inferentia (first generation): Four NeuronCores and 8 GB of DDR4 per chip, support for FP16, BF16, and INT8; Inf1 instances with up to 16 chips
  • Inferentia2: Two NeuronCores, 32 GB of HBM per chip, and 190 TFLOPS of FP16; Inf2 instances with up to twelve chips
  • Extended data types: Inferentia2 supports FP32, TF32, and configurable FP8 among other lower-precision formats
  • Distributed inference: Per AWS, Inf2 instances enable scale-out inference over ultra-high-speed connectivity between chips
  • AWS Neuron SDK: Native integration with PyTorch and TensorFlow, automatic casting of high-precision models to lower-precision formats
  • Framework support: PyTorch, TensorFlow, and MXNet
  • Efficiency: Per AWS, up to 50 percent better performance per watt for Inferentia2 versus comparable EC2 instances

Typical Use Cases

Large language models: AWS explicitly describes Inferentia2 as optimised for large language models and latent diffusion models.

Language and text processing: Natural language processing, language translation, text summarisation, and speech recognition.

Image and video generation: Generative image and video workloads in inference.

Personalisation and fraud detection: Recommendation systems and fraud detection with high request volumes.

Benefits

  • Up to 2.3x higher throughput and up to 70 percent lower cost per inference for Inf1 versus comparable GPU instances per the provider’s statement
  • Substantial gains in throughput and latency from Inferentia to Inferentia2 per AWS
  • Four times the memory capacity per chip for Inferentia2 compared with Inferentia
  • Up to 50 percent better performance per watt for Inferentia2 per the provider’s statement
  • Common frameworks run through the Neuron SDK without changing ecosystems
  • No separate line item; billed through the EC2 Inf instances

Integration with innFactory

As an AWS Reseller, innFactory supports you with AWS Inferentia: assessing your inference workloads, choosing between Inf1 and Inf2 instances, porting models via the AWS Neuron SDK, and measuring throughput, latency, and cost against your current environment.

Typical Use Cases

Inference for large language models
Image and video generation
Speech recognition
Personalisation and fraud detection

Technical Specifications

Generations Inferentia (EC2 Inf1), Inferentia2 (EC2 Inf2)
Sdk AWS Neuron SDK

Frequently Asked Questions

What is AWS Inferentia?

AWS Inferentia is a chip family designed by AWS for deep learning inference. The first generation powers the Amazon EC2 Inf1 instances, and Inferentia2 powers the Inf2 instances.

What specifications does AWS state for the first generation?

Per the product page, a first-generation Inferentia chip has four NeuronCores and 8 GB of DDR4 memory and supports the FP16, BF16, and INT8 data types. Inf1 instances offer up to 16 Inferentia chips. AWS states up to 2.3x higher throughput and up to 70 percent lower cost per inference than comparable GPU instances.

What distinguishes Inferentia2?

Per the product page, Inferentia2 delivers up to 4x higher throughput and up to 10x lower latency compared with Inferentia. Each chip has two NeuronCores, 32 GB of HBM memory, which is four times that of Inferentia, and 190 TFLOPS of FP16 performance. Supported data types include FP32, TF32, and configurable FP8 plus further lower-precision formats. AWS states up to 50 percent better performance per watt versus comparable EC2 instances.

How many chips do Inf2 instances offer?

Per the product page, Inf2 instances offer up to twelve Inferentia2 chips and enable scale-out distributed inference with ultra-high-speed connectivity between the chips.

How are Inferentia chips programmed?

Through the AWS Neuron SDK. Per the product page it integrates natively with popular ML frameworks such as PyTorch and TensorFlow, automatically casts high-precision models to lower-precision formats, and optimises performance. The page names PyTorch, TensorFlow, and MXNet as supported frameworks.

Which workloads suit Inferentia?

The product page names natural language processing, language translation, text summarisation, video and image generation, speech recognition, personalisation, and fraud detection. For Inferentia2, AWS highlights large language models and latent diffusion models.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of AWS (official documentation). This page does not represent an offer by AWS.

Quick Links

AWS Cloud Expertise

innFactory is an AWS Reseller with certified cloud architects. We provide consulting, implementation, and managed services for AWS.

Ready to start with AWS Inferentia: Chips for Deep Learning Inference?

Our certified AWS experts help you with architecture, integration, and optimization.

Schedule Consultation