overfeed.news

Share GPU clusters across teams with isolation and fairness using Amazon SageMaker HyperPod

1d

Age

Published
Collected
Image: AWS Machine Learning Blog

A reference architecture for securely sharing one Amazon SageMaker HyperPod EKS cluster across multiple teams, using AWS IAM Identity Center for authentication, per-team SageMaker Domains and Kubernetes namespaces for isolation, HyperPod Task Governance for fairness, and namespace-level cost allocation for chargeback.

Excerpt from the source

Multiple teams within the same company increasingly need shared access to expensive GPU clusters for their generative AI operations, while maintaining isolation boundaries, resource fairness, and operational independence. Consider a data science team training large language models, a computer vision group running inference workloads, and a research team experimenting with new model architectures. All of them might need access to the same cluster. Without a well-designed multi-tenant…

Read the full article at aws.amazon.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at AWS Machine Learning Blog.

Log in to follow this source

Introducing Claude Haiku 5.5 on AWS

Claude Haiku 5.5 is now available on Amazon Bedrock and Claude Platform on AWS. According to Anthropic, it is the fastest, most efficient model in the Claude 5.5 family, built for subagents and high-volume, cost-sensitive work, and costs around 75% less than Claude Haiku 4.5 for most tasks. This post covers its improvements and how to get started.

Share GPU clusters across teams with isolation and fairness…