Wrapture is a new indispensable monkey patching package for Python developers.
“This feels like one of those Swiss Army Knife packages that, once mastered, will provide value against all sorts of problems for years to come.”
71 tracked signals on observability.
Wrapture is a new indispensable monkey patching package for Python developers.
“This feels like one of those Swiss Army Knife packages that, once mastered, will provide value against all sorts of problems for years to come.”
Traditional golden signals are insufficient to monitor LLM app quality and safety
“a 200 OK response from your server doesn't necessarily mean that the response was useful or correct for your end user”
Weights & Biases released Agent Arya, a self-improving research agent using evaluation-driven self-reinforcement.
“tests, assessments, agents, and how you configure them are covariant”
Strands Evals and Amazon Bedrock AgentCore add skill-focused evaluators to measure agent skill selection and instruction following.
“A skill is a reusable set of instructions, usually stored in a SKILL.md file, that teaches an agent a domain-specific task like redacting a contract, reconciling an invoice, or following a team’s pull-request conventions.”
LangSmith Engine auto-monitors production agent traces and generates ready-to-merge code fixes
“Engine does have the time.”
Toyota's GearPal AI tool delivers six-figure annual savings per manufacturing line using LangChain.
“Enterprise AI will be on the balance sheet and it will be because of LangChain.”
LangChain launches tuned evaluators that beat frontier models on agent evals at lower cost
“our post-trained perceived error evaluator outperformed all frontier closed and open models, while still remaining the most cost-effective”
AI evals must evolve from LLM-as-judge to agent-as-judge as systems shift to multi-agent long-horizon tasks
“we didn't just make the problem harder, we actually got a fundamentally different type of problem”
AI agents lack persistent state checkpoints, making debugging and replay impossible today.
“all of that is lost and it is only stamped as a read-only trace by the end, which is sitting in another tool far away from where the actual code is.”
Production AI agents produce costly non-reproducible failures that pass silently with clean API responses, making them undebuggable.
“Setting the temperature to zero doesn't fix a broken reasoning path. It just means the model is going to make the exact same logical error, the exact same way, at the exact same time.”
Agentic AI evaluation must shift from model benchmarking to production-grade system behavior testing.
“The question is no longer did the model generate the right answer? The question is did the system behave correctly?”
AI engineering evals are fundamentally data science, and teams keep repeating avoidable eval mistakes.
“I'm saying that using LLM judges kind of blindly without validating your validators is bad.”
AWS released Agent-EvalKit, an open-source toolkit for systematically evaluating AI agents' full execution paths.
“An agent might deliver a well-structured, actionable response while hallucinating, fabricating facts because its tools returned empty results.”
Foundry Observability uses OpenTelemetry to unify tracing across any agent framework or cloud without rewriting agents.
“Today I'm going to show you how to answer those questions in one place, foundry observability, not by rewriting your agents into one agent framework, but by adopting open telemetry instrumentation with a few lines of code and without changing your existing agent logic.”
Microsoft Foundry observability adds end-to-end tracing, evaluation, monitoring, and ROI for AI agents on any framework.
“agents are non-deterministic, creating new reliability and consistency challenges for developers and operators”
Agent quality demands both pre-production evals and production observability to manage LLM non-determinism risks
“grading matters because LLMs are inherently non-deterministic”
Voice agent transcripts mask real failures that only audio-level observability reveals
“the logs lie”
AI agents and MCP are making traditional dashboards and query languages obsolete
“Every dashboard you've ever used, every weird little-known query language ever written, was a means of translation between you and your data, because the machines on the other end couldn't understand what you really wanted.”
LangChain has evolved from an open source framework into a full agent development lifecycle platform anchored by LangSmith.
“Turns out building the agent is really fun. Super fun. And kind of easy now, but actually keeping it from going completely sideways in production is the hard part.”
Salesforce touts Agentforce ROI playbook and launches Agent Optimizer with now-free, unmetered observability.
“But my favorite part of observability is that it's now unmetered. That means it's free, y'all.”
AWS launches observable enterprise agentic RAG using Bedrock Managed Knowledge Bases and AgentCore with MCP.
Observability and evaluations are the must-learn skills for the agentic AI era
“go learn about evaluations and evaluators. It's going to take you a long way.”
Observability and evaluations are the critical skills for the agentic AI era
“go learn about evaluations and evaluators. It's going to take you a long way.”
Microsoft Foundry offers end-to-end tooling for deploying, governing, and monitoring AI agents in production.
“The biggest challenge when building your AI agent is how do you get it in production?”
Clay runs 300M agent executions monthly and calls evals non-negotiable at that scale
“eval's became non-negotiable”
Podium uses LangSmith to trace agent chain-of-thought and debug unexpected AI behavior
“The reality is when you really get into the details of what context that agent was provided, It becomes obvious the agent was behaving rationally.”
Amazon OpenSearch Service MCP Apps embed interactive observability visualizations inside AI agent chat threads
“You're not trusting the AI's interpretation. You're seeing the actual query result rendered as an interactive chart, trace waterfall, or service map.”
AI agents evolve rapidly but evaluation frameworks fail to keep pace with model changes
“building a demo with AI is really easy but making it production quality is really hard”
AWS extends AgentCore Observability to monitor AI agents on any cloud or on-premises
LangChain frames agent improvement as a data mining problem over production trace data
“there's a very tight coupling between what observability is and what continual learning is”
70% of engineer time goes to running production, not writing code, and AI is making this worse.
“AI is creating a lot more issues in production as, you know, AI code sort of goes through.”
Madrigal Pharmaceuticals cut AI agent time-to-production from 12 weeks to 2 using LangChain
“The first time I looked at a trace in LangSmith, I felt like I was peering into the brain of our AI agent.”
LangSmith now provides full observability for voice agents built on Gemini Live
“To take that agent to production safely, you need visibility into what your agent is doing.”
Arize founder argues observability is evolving from human dashboards to autonomous self-healing agent systems
“how do I build systems that autonomously fix themselves?”
Traversal builds SRE agents that search petabyte-scale logs with no labeled training data
“a single investigation's probably gonna cost you the GDP of a small country”
AI self-improvement loops require domain-specific evaluation functions, not just agentic design.
“does the code compile or not”
Amazon Bedrock AgentCore Observability gives layered visibility into AI agent execution to debug silent production failures.
“Production artificial intelligence (AI) agents can fail silently.”
Local on-device models can replace frontier models like GPT-5 and Claude to cut inference costs, latency, and security risks.
“Every time you reach for foundation models like GPT-5 or Claude, it's costing you, your users, and the environment.”
Agent failures stem from static retrieval that never learns from eval and observability signals.
“We made wrong answers appear faster and cheaper, but we forgot to make retrieval learn.”
OpenGov built and scaled OG Assist, a production AI agent embedded across its government ERP product suite.
AWS demoed Bedrock AgentCore features enabling self-improving agents via evaluations, observability, and prompt optimization.
“I drive obserability evaluations, behavioral analysis and recommendation set of capabilities for agents.”
AWS's Strands Evals SDK adds detectors that automatically diagnose AI agent failures and recommend fixes.
“Detectors answer “why did it fail?” by producing diagnoses at the per-span level with categorized failures, causal chains, and fix recommendations.”
PostHog is building a pipeline that turns product signals into auto-generated pull requests via background agents.
“instead of ever looking at your analytics dashboard, or your errors, or your logs, we just want you to look at PRs that are ready for you in GitHub”
Apple's Instruments now profiles Foundation Models apps, giving developers observability into on-device and server LLM agent loops.
“Traditional code is predictable. LLMs are non-deterministic - the same input can produce different outputs.”
Most developers building agents lack visibility into what their agent does under the hood; observability and evals close that gap.
“you probably don't know what your agent is actually doing under the hood”
Traditional observability breaks when agents, not humans, become the primary consumers of telemetry.
“Logs were built for events. Agents produce decision breadcrumbs.”
GenAI applications require extending traditional golden signals because they are non-deterministic, variably costed, newly attackable, and subjectively judged.
“What makes Genai applications fundamentally different from the traditional software we've been monitoring for decades? There are four key shifts.”
AWS introduces AgentOps practices on Bedrock AgentCore to operationalize agentic AI in production.
“That's where AgentOps comes in, the operational discipline for deploying, managing, and continuously improving AI agents in production.”
LangSmith on AWS provides five evaluation patterns for deep agents in production
“Validating AI agent behavior before production is one of the hardest problems in applied AI.”
Google recommends five architectural patterns for production-ready AI agents at scale
“Production agents need robust architecture, not just clever prompts.”
Zip reduced AI feature development from weeks to days using LangGraph and LangSmith
“if you spend time developing many components yourself, you are falling behind because you are not keeping up with these rapid changes”
AI coding agents accelerate development but are shifting engineer time toward troubleshooting, not design.
“more and more time is spent on troubleshooting. Much more code is being written. People are less understanding of the code that goes into production.”
Wonder uses LangSmith traces and LLM-as-judge to automate AI agent evaluation at production scale
“The feedback loop literally went from minutes, sometimes even like multiple hours, versus now it's all automated.”
Morningstar deployed LangSmith to gain observability into production AI agents after flying blind
“Without LangSmith, we didn't have a common language for observability and tracing. And so debugging becomes anecdotal.”
AWS enables native Codex usage observability via OpenTelemetry and CloudWatch without a centralized proxy.
“It is no longer only, 'Can this tool help a developer?' It becomes, 'How do we understand adoption, manage consumption, maintain reliability, and scale access responsibly?'”
Google ADK voice agents using Gemini Live can be traced with LangSmith for production observability
“To take that agent to production safely, you need visibility into what your agent is doing and the ability to test its behavior in a number of different scenarios that it might encounter with real end users.”
LangSmith now traces every Cursor agent turn with full tool call and sub-agent visibility
“If you watch the Claude Code or Codex versions of the series, you'll notice that the traces actually share a common structure under the hood, so you can compare all three agents in the same LangSmith workspace.”
AWS shows how to build an agentic incident triage assistant using Amazon Quick, New Relic MCP Server, and Asana.
“From a single prompt, the Amazon Quick agent investigates the incident, assembles a root cause analysis (RCA) brief with evidence links, and creates a tracked Asana task ready for handoff.”
Building reliable AI systems requires observability, evaluation, and experimentation, not magic.
“It's really the same set of patterns, just maybe a different flavor coming out. And it's really it feels like magic, but it's not magic, right? It's all just engineering.”
Microsoft Foundry's agent platform unifies tracing, evaluation, and optimization to test non-deterministic AI agents end-to-end.
“The properties that makes agents so useful like they're stateful, they they're long horizon, they can plan, they can course correct, they interact with their environment through tools are also the ones that makes it very hard to test.”