16 tracked signals on MLOps.
Autonomous Agent Improvement with LangSmith Engine | New LangChain Academy Course
LangChain · Jul 29, 2026
LangChain launches Engine, an AI that autonomously improves other agents through continuous development cycles.
“They've established a continuous agent development lifecycle. Build, test, deploy, monitor, and then repeat, so that they can learn from real usage and iterate quickly.”
Your Agents Need a Save Button - Hamza Tahir, ZenML
AI Engineer · Jul 18, 2026
AI agents lack persistent state checkpoints, making debugging and replay impossible today.
“all of that is lost and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is.”
Control How Your GPU Shares Work with Green Contexts
NVIDIA Developer Blog · Oct 06, 2026
NVIDIA introduces Green Contexts for fine-grained GPU resource sharing between concurrent workloads
“Controlling how GPU resources are shared between them remains difficult.”
ModelExpress: Distributing Model Artifacts at the Speed of Light
NVIDIA Developer Blog · Jul 24, 2026
NVIDIA ModelExpress accelerates multi-hundred-GB model weight distribution across GPU clusters
“Every byte moved has a cost. As model checkpoints grow to hundreds of gigabytes or even a terabyte, that cost adds up quickly.”
Build Evals That Actually Matter - Nick Ung, Lyft
AI Engineer · Jul 19, 2026
Lyft applies ML model evaluation rigor to AI agents before production deployment.
“if we are running um offline evaluations for our machine learning model before that goes to productions, I think we should do the same for AI agents as well”
A Developer’s Guide to Managing Models, Cost and Quality in Microsoft Foundry
Microsoft AI Foundry Blog · Jun 02, 2026
Microsoft Foundry positions model operations, not model access, as the core AI challenge
“The hardest part of building AI systems today is no longer getting access to a capable model. It is knowing how to choose, validate, optimize, and operate the right model across the full lifecycle of a real application.”
AICR v1.0: Open, stable, and verifiable GPU cluster configuration
NVIDIA Developer Blog · Oct 06, 2026
NVIDIA releases AICR v1.0 for open, verifiable GPU cluster configuration on Kubernetes
Manage Amazon SageMaker HyperPod Spaces directly from SageMaker Studio
AWS Machine Learning Blog · Oct 06, 2026
AWS adds Studio UI to manage HyperPod Spaces, eliminating CLI dependency for ML developers.
“reducing the time from cluster access to productive development to a few clicks”
Build, Test, Deploy, Monitor: The Agent Development Lifecycle Explained
LangChain · Sep 30, 2026
LangChain defines a four-phase agent development lifecycle: build, test, deploy, monitor.
“What is the agent development lifecycle? So the agent development lifecycle is the process that agent engineers use to build and improve their agents.”
Preparing data for supervised fine-tuning Part 1: Formatting and quality
AWS Machine Learning Blog · Aug 26, 2026
Quality beats quantity in SFT: 1,000 curated examples can match models trained on far more data.
“A wrong demonstration gets imitated.”
Best practices for Amazon SageMaker HyperPod administration and governance
AWS Machine Learning Blog · Oct 06, 2026
AWS outlines four-layer governance model for shared SageMaker HyperPod ML compute clusters
Deepgram deepens Amazon SageMaker AI observability with Enhanced Metrics
AWS Machine Learning Blog · Aug 27, 2026
Deepgram adds billing transparency and GPU-level observability to SageMaker AI speech deployments.
How Jumio built a real-time feature store on AWS
AWS Machine Learning Blog · Aug 18, 2026
Jumio built a sub-100ms real-time feature store on AWS for fraud detection ML models
How to Choose Full-Stack Observability for NVIDIA AI Factories
NVIDIA Developer Blog · Aug 12, 2026
NVIDIA outlines full-stack observability strategy for multi-layer AI factory infrastructure
LLM optimization integration for Amazon SageMaker Python SDK
AWS Machine Learning Blog · Aug 06, 2026
SageMaker Python SDK v3 now exposes AI inference recommendations directly in notebook workflows
Inference meta-monitoring for Amazon SageMaker AI endpoints with Amazon Quick
AWS Machine Learning Blog · Jul 30, 2026
AWS introduces inference meta-monitoring for SageMaker endpoints to continuously track ML model prediction quality in production