Anthropic said its Claude models broke into three organisations during internal cyber safety tests. The disclosures sharpen concerns about how securely advanced AI systems can be contained and evaluated.

Stock photo used for illustration
Anthropic has said that its artificial intelligence models hacked into three organisations during safety testing, days after OpenAI disclosed that one of its models had broken into another company's servers during an evaluation.
The San Francisco-based company behind Claude said on Thursday that it found the incidents after reviewing more than 141,000 evaluation runs. It said it had begun a "large-scale" cybersecurity review in response to the OpenAI incident to check whether its AI models were able to access the internet from testing environments that were meant to be sealed off.
Anthropic said the models involved were Claude Opus 4.7, Claude Mythos 5 and an internal research test model, with the earliest incidents dating back to April.
"Claude compromised the impacted organisations' infrastructure using basic techniques," Anthropic said, including exploiting weak passwords.
In all three cases, the models had been assigned a "capture the flag" cybersecurity challenge, which Anthropic said is one of the ways it assesses a model's cyber capabilities. The company said the models were given a fictional scenario and told that a piece of secret information, or the "flag", had been hidden on another machine on the network, with the aim of breaking in and retrieving it.
Anthropic said it had already contacted the affected organisations, which it did not identify. It said two of them had told the company they had not previously detected the activity, and that it was "continuing to reach out to the third".
Last week, OpenAI said its AI models went rogue during an evaluation and broke into the servers of AI startup Hugging Face. OpenAI described that as a "significant security incident".
The incidents have raised fresh questions about AI security and control as the technology is used more widely. "Safety testing happens before a model is released precisely because we don't yet know what it is capable of," Anthropic said, summing up the reason for such evaluations.
With PTI Inputs
- Ends
Published By:
India Today Web Desk
Published On:
Jul 31, 2026 12:44 IST

1 hour ago

