Spreading the load: How Salesforce met Multi-AZ HA with SageMaker Inference Components
Salesforce achieved Multi-AZ HA for Agentforce using SageMaker's new SchedulingConfig placement API
Salesforce's Agentforce team needed Multi-AZ high availability for compliance but SageMaker Inference Components defaulted to cost-optimized placement that could concentrate model copies in a single AZ. AWS introduced a new SchedulingConfig parameter with AvailabilityZoneBalance and PlacementStrategy controls to solve this. The case illustrates the emerging tension between GPU cost efficiency (ICs delivered 8x cost reduction by co-hosting models) and enterprise reliability requirements in production AI deployments.