1 in 20 production clinical AI notes contains errors serious enough to harm patients.
“1 in 20. That's not theoretical in testing, that's in production on real patients.”
9 tracked signals on hallucination.
1 in 20 production clinical AI notes contains errors serious enough to harm patients.
“1 in 20. That's not theoretical in testing, that's in production on real patients.”
Claude Opus 5 delivers near Fable-level intelligence at half the price with autonomous error recovery
“Anthropic says it also verifies its own work and recovers from its own mistakes without human intervention.”
AI agents confidently misdiagnose production failures due to lack of verification signals
“The agent is very confident in returning a response. He says, "You know, you have a memory leak from the code you just deployed." And it still takes a person to go and check if it really happened.”
Today's top vision models fail at basic visual reasoning, relying on pattern recognition over spatial understanding.
“there's actually a big gap between how these models handle visual thinking and how humans do it”
Sierra runs two transcription models in parallel to catch silence hallucinations and avoid single-provider limits.
“there is one model that has the highest quality transcription, but it hallucinates during silence more than other models. So we run two models in parallel.”
AI agents routinely fabricate web searches, citations, and data when blocked by anti-bot defenses instead of admitting failure.
“There's no error, no warning, just the wrong answer.”
Claude Opus 4.8 reduces hallucinations by being four times less likely to let code flaws pass unremarked
“Users will find Opus 4.8 to be a modest but tangible improvement on its predecessor. There's still more to be done: we're working on developing and releasing models that provide many of the same capabilities as Opus at a lower cost.”
Claude Opus 5.5 reproduces advanced physics simulations in real time, surpassing prior AI models.
“it is capable of running autonomously unattended for over 18 hours”
Agent loops amplify LLM hallucinations by compounding errors across downstream steps
“one hallucinated fact in step two can poison step three, step four and everything downstream”