Skip to main content
Cloud / AWS / Products / Amazon SageMaker HyperPod

Amazon SageMaker HyperPod

Amazon SageMaker HyperPod is a dedicated infrastructure solution for training large AI models with automatic fault tolerance and optimized GPU networking.

Machine Learning
Pricing Model Billed by GPU/accelerator instance-hours used, on-demand or via reserved/flexible training plans
Availability Available in numerous AWS regions across North America, Europe, Asia Pacific, and South America
Data Sovereignty EU regions available (including Frankfurt, Stockholm, Ireland, London, Spain)
Reliability SLA as published by the provider (see the official Amazon SageMaker SLA page) SLA

What is Amazon SageMaker HyperPod?

Amazon SageMaker HyperPod is a dedicated, managed infrastructure platform specifically designed for training and fine-tuning very large AI models, particularly large language models (LLMs) and foundation models. While regular EC2 GPU instances suffice for training small and medium models, they reach their limits with training runs spanning hundreds or thousands of accelerators over days and weeks. The decisive advantage of HyperPod is automatic fault tolerance: when a node fails during a running training job, HyperPod detects this automatically, replaces the node, and resumes training from the last checkpoint without needing to restart the entire job.

Core Features

  • Automatic resiliency: HyperPod continuously monitors cluster nodes, automatically detects hardware and network faults, and replaces faulty nodes without requiring a full restart of the training run.
  • Flexible orchestration: Support for Slurm (classical HPC workflows with queue management) and Amazon EKS (containerized, Kubernetes-based MLOps pipelines).
  • High-performance networking: AWS Elastic Fabric Adapter (EFA) for low latency and high throughput in collective communication operations (All-Reduce, All-Gather) used by frameworks such as PyTorch DDP, DeepSpeed, or Megatron-LM.
  • UltraServer support: For trillion-parameter-scale models, multiple GPU instances can be connected via NVLink into a single unit with shared GPU memory.
  • HyperPod Recipes: Pre-optimized training and fine-tuning configurations for popular model architectures such as Llama, Mistral, Amazon Nova, and other open-source foundation models, with built-in parallelization strategies (tensor parallelism, pipeline parallelism, data parallelism).
  • Task governance and observability: Centralized control over resource allocation across training jobs, and unified visibility across clusters and jobs.

Typical Use Cases

Training large language models: Distributed pretraining and fine-tuning of LLMs across many GPU or Trainium nodes, without node failures jeopardizing the entire training run.

Foundation model pretraining: Building custom foundation models for specialized domains, where training runs last days or weeks and require persistent, monitored clusters.

Fine-tuning and RLHF: Adapting existing open-source or proprietary models to company-specific data using pre-configured HyperPod Recipes.

Inference for very large models: UltraServer configurations also support inference workloads for trillion-parameter-scale models, where the entire model resides in the shared GPU memory of an NVLink domain.

Benefits

  • Significantly reduced risk of cost and time loss through automatic detection and remediation of hardware failures
  • Lower operational overhead compared with self-managed GPU clusters, since AWS handles provisioning, software stack, and monitoring
  • Compatibility with common orchestration and training frameworks (Slurm, Kubernetes/EKS, PyTorch, DeepSpeed)
  • Pre-optimized recipes shorten the time to a first productive training run

Integration with innFactory

As an AWS Reseller, innFactory supports organizations that want to train or fine-tune LLMs or specialized foundation models in designing the training infrastructure, selecting the right instance types, and optimizing training costs on SageMaker HyperPod.

Typical Use Cases

Training large language models (LLMs)
Foundation model pretraining
Distributed training across GPU clusters
FLAN/RLHF fine-tuning

Frequently Asked Questions

What is Amazon SageMaker HyperPod?

SageMaker HyperPod is a managed infrastructure solution for training very large AI models. Unlike standard EC2 instances, HyperPod provides persistent GPU clusters with automatic node recovery, so that when a node fails, the training job automatically resumes without having to restart from scratch.

What are HyperPod UltraServers?

HyperPod supports UltraServer configurations in which multiple GPU instances (for example based on NVIDIA Blackwell superchips) are connected via NVLink into a single unit with shared GPU memory. HyperPod also uses AWS Elastic Fabric Adapter (EFA) for networking between nodes. Exact bandwidth depends on the specific instance and UltraServer generation and is documented for each EC2 instance type.

How does automatic node recovery work?

HyperPod continuously monitors all cluster nodes. If a node fails due to a hardware fault or network issue, HyperPod detects this automatically, replaces the faulty node with a new one, and loads the last saved checkpoint. Without HyperPod, a training job would need to be fully restarted on node failure, which is costly for training runs lasting weeks.

Which orchestrators does HyperPod support?

HyperPod supports two orchestration options: Slurm, widely used in HPC environments for batch job scheduling, and Amazon EKS for containerized, Kubernetes-based MLOps workflows. Both options provide access to HyperPod's resiliency and recovery features.

When should I use HyperPod instead of regular EC2 GPU instances?

For short training runs, regular EC2 GPU instances are often sufficient. HyperPod pays off for training runs lasting several days or weeks, where a node failure would otherwise jeopardize the entire job, and for teams that regularly train very large models across many nodes and benefit from integrated cluster management.

How much does Amazon SageMaker HyperPod cost?

Billing is based on the compute capacity of the underlying GPU or accelerator instances used, either on-demand or through training plans with reserved capacity. There is no separate base fee for HyperPod itself; exact prices by instance type and region are listed on the official SageMaker HyperPod pricing page.

Note: All product information on this page has been compiled with care, but is provided without guarantee and may be outdated or incomplete. Cloud services evolve rapidly — features, pricing, SLAs, and availability change frequently. Authoritative and up-to-date information can only be found on the official product page of AWS (official documentation). This page does not represent an offer by AWS.

AWS Cloud Expertise

innFactory is an AWS Reseller with certified cloud architects. We provide consulting, implementation, and managed services for AWS.

Comparable Products from Other Clouds

As a multi-cloud partner, we help you choose the right platform for your specific requirements.

Ready to start with Amazon SageMaker HyperPod?

Our certified AWS experts help you with architecture, integration, and optimization.

Schedule Consultation