The Hallway Track
Product Launches

Reduce inference cold starts on Amazon SageMaker HyperPod with model caching

AWS Machine Learning Blog · Sep 10, 2026 · Product Launches

Amazon SageMaker HyperPod introduces model caching to reduce inference cold start times.

“your pods can typically start serving traffic in seconds rather than tens of minutes.”

AWS has launched model caching for Amazon SageMaker HyperPod, which allows faster inference by pre-loading model weights and container images on local storage. This significantly reduces the cold start time from minutes to seconds, enhancing scalability and performance during traffic spikes.

AWS model_caching SageMaker

Watch / read the original source →