LemonSlice launched a Microsoft-partnered real-time Teddy Roosevelt avatar demonstrated to President Trump
“What we mean by this is making an Avatar that is indistinguishable from a human on a video call.”
14 tracked signals on voice-agents.
LemonSlice launched a Microsoft-partnered real-time Teddy Roosevelt avatar demonstrated to President Trump
“What we mean by this is making an Avatar that is indistinguishable from a human on a video call.”
Satya Nadella frames agent architecture as a harness looping across models, data, and tools.
“you kind of want to harness to define the models, the the data, uh and the tools. And so that you have a loop across those three.”
Voice agent transcripts mask real failures that only audio-level observability reveals
“the logs lie”
LangSmith now provides full observability for voice agents built on Gemini Live
“To take that agent to production safely, you need visibility into what your agent is doing.”
Hugging Face benchmarks frontier ASR systems on code-switched bilingual speech for voice agents.
AWS released the open source Nova Sonic Test Harness to automatically evaluate voice agents at scale without a microphone.
“It runs complete multi-turn conversations with Amazon Nova Sonic automatically, evaluates them using LLM-as-judge techniques, and can even detect cases where the model's audio output doesn't match its text output (audio hallucinations).”
Real-time multimodal voice agents are bottlenecked by media infrastructure, not models, which LiveKit handles via WebRTC on Azure.
“You see models they are the easy part. Now everything around it is the hard part.”
Natera's healthcare voice agent hit 100% tool-calling accuracy at under $0.01 per call using Bedrock AgentCore
Google ADK voice agents using Gemini Live can be traced with LangSmith for production observability
“To take that agent to production safely, you need visibility into what your agent is doing and the ability to test its behavior in a number of different scenarios that it might encounter with real end users.”
AWS shows how to build a healthcare voice appointment agent using Amazon Nova 2 Sonic and Bedrock AgentCore.
“Instead of chaining separate transcription, reasoning, and synthesis services, Nova 2 Sonic processes speech natively in a single model—so vocal context like tone and pace isn’t lost to transcription.”
LangChain shows how to convert an existing LangGraph agent into a voice agent using the Pipecat framework.
“Pipecat is going to handle all of the glue required to take the input audio, convert it to text, run it through our LangGraph LLM layer, then convert that back to speech in order to send it back to the end user.”
DeepLearning.AI launched a course on building fast, reliable voice agents in partnership with Vocal Bridge.
“Voice is an under-exploited frontier.”
AWS solutions architects describe building an enterprise voice AI agent for healthcare using Amazon Bedrock AgentCore.
“we recently work with a healthcare customer to help them build out their agentic platform using the voice agent as a use case”
Microsoft demos building a multimodal voice agent that books flights and hotels conversationally.
“lately almost every conversation comes back to one thing, agents”