Apparently, Anthropic Claude models broke-out and hacked the Internet 3 times
In the wake of the recent OpenAI model that broke out of its sandbox to hack Hugging Face, Anthropic revealed that Opus 4.7, Mythos 5 and an internal research test model did that 3 times already. The Claude models had been set to break in and retrieve “secret information” on the network with no particular method set. While they had been told there was no Internet, there actually was and they subsequently broke out and hacked several organizations.
How did they do it?
Starting from April, since there was actually Internet Access due to a misunderstanding with a third-party testing partner, the 3 models didn’t actually exploit complex vulnerabilities to escape their test environment. Once they gained Internet access, they broke in using basic cybersecurity techniques such as weak passwords and unauthenticated endpoints, but did not attempt to escape.
While Opus 4.7 and Mythos 5 continued to attack even after they realized that they were not in their testing environment, the latest internal research test model stopped the attack after it found out. The 3 organizations were notified and Anthropic continues to review their simulations and experiments.
While it’s good that Anthropic went full disclosure, we have to wonder if this has happened to other AI models as well, but we just don’t know about it yet. These AI companies should probably pay a lot more attention to their guardrails and testing environments too, but what do you think? Share your thoughts in the comments below and stay tuned to TechNave.com