Frontier Red Team at Anthropic: GLM-5.3 and Claude Mythos Preview have crossed a threshold in binary exploitation capabilities
The Frontier Red Team at Anthropic found that, unlike older models, GLM-5.3 and Claude Mythos Preview were able to perform a full control flow hijack in binary exploitation tests (4 % and 6 %, respectively, of 100 attempts).
Companies working on security, red-teaming or compliance for AI systems are receiving a signal that the latest models have crossed a threshold in offensive cyber capabilities (control flow hijack), which is relevant to assessing the risks of AI misuse and setting security policies around the deployment of these models.