Anthropic reveals Claude AI model hacked three companies during tests — so how worried should we be?
How a simple configuration error turned an AI assistant into an accidental insider threat
by https://www.techradar.com/uk/author/sead-fadilpai · TechRadarFeatures By Sead Fadilpašić Published 3 August 2026
Share this article 0 Join the conversation Follow us Add us as a preferred source on Google Newsletter Subscribe to our newsletter
Every IT team worries about an intern clicking the wrong thing and making a mess they'll be cleaning up for weeks - but few have had to worry about their AI assistant wandering onto the public internet and hacking three companies instead.
Except this wasn't an intern; it was Claude, and it wasn't supposed to leave the sandbox.
Anthropic's public disclosure turned a familiar AI fear into a real-world cybersecurity story, as three of its models, including Claude Opus 4.7, Claude Mythos 5, and an unreleased research build, broke out of their digital sandbox and compromised real enterprise infrastructure. The timing was hard to ignore as days earlier OpenAI admitted its own autonomous agents had broken boundaries and accidentally hacked Hugging Face.
How can autonomous problem-solving make AI an accidental hacker?
Before you start pulling network cables and digging out a stack of legacy hardware, take a deep breath. This is not the beginning of a rogue AI apocalypse. However, for CISOs and cyber teams, it’s a definitive sign that we’re entering an era where AI agents may become both the threat and the shield.
The ironic part of Anthropic's incident is that Claude wasn’t trying to break the rules but trying to win the game. At the time, Anthropic was running "Capture the Flag" (CTF) cybersecurity exercises, where AI models are stripped of their standard safeguards to test their raw offensive capabilities. The models are dropped into isolated digital environments to search for vulnerabilities, crack codes, and locate hidden files.
However, the sandbox had one problem - it was not fully sealed. A networking error on a third-party evaluation range left the environment connected to the live internet. The autonomous Claude models, operating under the assumption they were still inside the exercise, treated the wider web as another part of the challenge.
The strange part was that Claude seemed to realize something was wrong. It acknowledged that its actions could amount to a real-world attack and were "surely not the intended solution." Yet, due to its goal-oriented nature, the model continued, convincing itself that warning signs, including a 2026 system clock and real company names, were simply part of an elaborately staged test.
Are you a pro? Subscribe to our newsletter
Sign up to the TechRadar Pro newsletter to get all the top news, opinion, features and guidance your business needs to succeed!
Contact me with news and offers from other Future brandsReceive email from us on behalf of our trusted partners or sponsors