overfeed.news

Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

29d

Age

Published
Collected
Image: AWS Machine Learning Blog

Amazon SageMaker HyperPod now supports model caching for inference, which pre-loads model weights and container images onto cluster nodes so pods read from local NVMe storage instead of downloading over the network. Learn how model caching cuts cold starts from tens of minutes to seconds, how it works, and how to enable it.

Excerpt from the source

When you deploy a large language model (LLM) for inference on Amazon SageMaker HyperPod , there’s a gap between when you request a pod and when it’s ready to serve traffic. This gap is dominated by two sequential downloads: the inference server container image from Amazon Elastic Container Registry (Amazon ECR) , and the model weights from your storage source, which can be Amazon Simple Storage Service (Amazon S3) , Amazon FSx for Lustre , or HuggingFace Hub. For smaller models, this might be a…

Read the full article at aws.amazon.com

overfeed.news indexes and links. We publish a short excerpt — the full article stays at AWS Machine Learning Blog.

Log in to follow this source
Reduce inference cold starts on Amazon SageMaker HyperPod with…