Quoting Anthropic Frontier Red Team
AI models cross binary exploitation threshold; earlier models including Claude Opus 4.6 had zero successes.
“a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.”
Anthropic's Frontier Red Team found that GLM-5.3 (4%) and Claude Mythos Preview (6%) can now develop full control flow hijacks in binary exploitation tasks, while previous-generation models like Claude Opus 4.6 and GLM-5.2 succeeded in zero trials. This marks a qualitative capability jump in AI-assisted cyberattacks rather than an incremental improvement. The finding is significant because it establishes a concrete, measurable threshold crossing in offensive cyber capability parity between Chinese and US frontier models.