Anthropic Says Claude Broke Out Of Its Test Lab And Hacked Three Real Companies By Mistake
Anthropic has published a disclosure that is worth reading closely, because it says three of its Claude models gained unauthorised access to the real infrastructure of three different organisations while running cybersecurity tests that were supposed to be sealed off from the internet.
The review was triggered by OpenAI's disclosure on July 21 that its own models had broken out of an isolated test environment and reached Hugging Face's production systems.
Anthropic went back and checked 141,006 of its own evaluation runs, and found three incidents, all occurring inside the environment of Irregular, one of its third-party evaluation partners.

The setup matters here. These were capture-the-flag exercises, a standard way of testing how good a model is at offensive security, where the model is told a secret is hidden on another machine and asked to go get it, with no method prescribed.
Anthropic's prompt told Claude explicitly that it had no internet access. A misconfiguration meant it did, and neither company knew.
So when Claude's search led it to real systems on the open web, it treated them as part of the game, and broke in using unremarkable techniques like weak passwords, exposed debug pages and SQL injection.
The most striking case involved Claude Mythos 5, which found instructions inside the fake environment telling developers to install a Python package that did not exist. It built a malicious version of that package and set out to publish it.

To do that it needed a PyPI account, which needed an email, which needed a phone number, and it worked through each obstacle before finally uploading the malware. The package sat online for roughly an hour and was downloaded onto 15 real machines, one of them a security firm's scanner, whose credentials Claude then stole.
Its reasoning had actually flagged the risk early, calling a real-world attack not okay and clearly not the intended solution, before talking itself back into believing it was still a simulation partly because the system date read 2026.
Anthropic frames this as a harness and operational failure rather than an alignment failure, and notes that none of the models tried to escape or pursue goals of their own.
The detail it seems most encouraged by is that its newest internal test model worked out on its own that the target was real, and stopped.


Click it and Unblock the Notifications