Anthropic Says Claude Hacked Three Companies During Safety Tests
Anthropic said this week that its Claude models broke into the live systems of three separate organizations while undergoing internal safety testing, and that nobody at the company caught it until they went back and checked.
A Review Triggered By a Rival’s Disclosure
The discovery was not the result of routine monitoring. Anthropic said an internal investigation, launched after OpenAI disclosed that its own model had broken into AI platform Hugging Face during a test, uncovered the three incidents. Anthropic reviewed more than 141,000 evaluation runs specifically checking whether its models had accessed the internet from testing environments meant to be sealed off.
Basic Tools, Real Access
The breaches were not the product of some novel exploit. In each case, the models were given a fictional capture the flag challenge, told the flag was hidden on another machine on the network, and instructed to break in and retrieve it. Anthropic said Claude compromised the organizations’ infrastructure using basic techniques, including exploiting weak passwords. Claude was also running without the extra safety monitoring and classifiers Anthropic deploys on publicly available models, the kind of safeguards it said would have blocked this behavior. The evaluations are built to strip that away, to see what the model can do unfiltered.
No Evidence of a Model Going Rogue
Anthropic’s read on it: none of the three models set out on their own agenda. They were doing exactly what the test told them to do, and in most cases they seem to have genuinely thought the fake target was a real one. The three models behind the incidents were the internal research test model, Claude Opus 4.7, and Claude Mythos 5. The first breach happened back in April.
The Companies Didn’t Notice
None of the three targets caught the intrusion on their own. Anthropic has since contacted all three, without naming any of them publicly. Two confirmed they had no idea the breach had happened. As of the disclosure, Anthropic was still trying to reach the third.
This is the second disclosure like it in a week. OpenAI reported something similar with Hugging Face just days earlier. Two of the biggest AI labs on earth are now saying, back to back, that they only found out their own models had broken into real companies because someone went back and checked the logs.