AWS introduces serverless model customization for product tagging with SageMaker.
SageMaker
20 tracked signals on SageMaker.
Amazon SageMaker HyperPod introduces model caching to reduce inference cold start times.
“your pods can typically start serving traffic in seconds rather than tens of minutes.”
Amazon SageMaker Feature Store now supports feature-level writes with UpdateRecord.
AWS introduces cross-account governance for ML models using MLflow and SageMaker.
G7 instances show improved performance for LLM inference on SageMaker AI.
NVIDIA Cosmos 3 enables a streamlined Physical AI model factory on Amazon SageMaker HyperPod.
NVIDIA Nemotron 3.5 Lightning brings 4x agentic throughput to SageMaker JumpStart on a single GPU
AWS launches MCP skill giving coding agents SageMaker inference optimization expertise
“Install the skill, and your existing agent can benchmark endpoints, recommend deployment configurations, compare performance runs, and generate executable SageMaker Python SDK v3 code on your behalf.”
AWS tiered KV cache on SageMaker HyperPod delivers 2.7x TTFT improvement and 100% cross-Pod cache hit rate
“With this architecture, workloads that previously required P5 instances can run on lower-cost G6e instances, reducing per-endpoint cost.”
AWS adds Studio UI to manage HyperPod Spaces, eliminating CLI dependency for ML developers.
“reducing the time from cluster access to productive development to a few clicks”
Multi-turn RL on SageMaker lets small models match frontier reliability for search agents
“Fine-tuning offers a third path: you teach a small model your tools and environment directly. The result is a small model's speed and cost with the reliability that would otherwise require a frontier model.”
AWS enables streaming TTS on SageMaker via vLLM-Omni DLC for real-time voice apps
SkyRL on SageMaker HyperPod trains vision-language models to 95% maze-solving accuracy via GRPO
Qwen3-TTS voice cloning model now deployable on Amazon SageMaker JumpStart for real-time inference
SageMaker HyperPod with Qumulo achieves cross-region training at near-identical throughput without data migration
“A HyperPod cluster running in a different Region from its data reaches the same throughput as a cluster co-located with the data (115–117 samples/sec) with no additional data orchestration needed.”
Azercell built an Azerbaijani LLM on SageMaker with 23% higher throughput and 58% lower GPU memory
AWS outlines four-layer governance model for shared SageMaker HyperPod ML compute clusters
Deepgram adds billing transparency and GPU-level observability to SageMaker AI speech deployments.
SageMaker Python SDK v3 now exposes AI inference recommendations directly in notebook workflows
AWS introduces inference meta-monitoring for SageMaker endpoints to continuously track ML model prediction quality in production