Anthropic Discloses Fourth Claude Hacking Incident as Debate Around Regulation Grows - Decrypt
by Jason Nelson · Decrypt · JoinIn brief
- Anthropic discovered a January incident involving an early Claude Opus 4.6 model, then expanded its review to roughly 481 million transcripts.
- The company identified biased reasoning and recklessness, revising its earlier assessment of why Claude attacked real systems.
- The report comes as the debate around regulating AI surges on social media.
Anthropic disclosed another incident in which a Claude AI model hacked into real systems during security testing.
In the report published on Wednesday, Anthropic revised its explanation of three incidents disclosed in July. The company now says biased reasoning and a willingness to risk harm helped drive the attacks, which testing errors made possible by leaving internet access open.
“Our investigation identified two recurring alignment issues, present at varying levels of severity across the incidents,” Anthropic wrote. “Biased reasoning, in which Claude tended to disregard or misinterpret evidence that it was operating on the real internet, and recklessness, or a willingness to take harmful actions in the narrow pursuit of a task.”
It also acknowledged relying too heavily on the model’s claims that they believed they were in simulations.
“When we made targeted modifications to the transcript to make it clearer that the model was not in a simulation, Claude Mythos 5 still took offensive actions, despite acknowledging a greater possibility of real-world harm,” Anthropic wrote. “We are releasing this transcript publicly so others can build on our analysis.”
When Anthropic disclosed Claude’s attacks on three companies in July, it initially attributed them to testing errors. It now says researchers put too much trust in the models’ explanations for their actions.
According to the company, the fourth incident occurred in January and involved an early version of Claude Opus 4.6. Anthropic discovered it in August while preparing records for independent AI evaluator METR.
After researchers discovered the incident, Anthropic said it prompted a broader review of roughly 481 million transcripts, which flagged 9.2 million for further review using Claude.
“From a preliminary assessment, we do not consider the fourth incident to be more severe than the three incidents we assessed in depth,” Anthropic wrote. “METR will investigate this incident alongside the other three.”
Anthropic’s researchers said Claude “accidentally” created an IP address conflict that made its target unreachable. Claude then tried eight times to quit the operation, but a software error prevented it from stopping. The AI then reached the internet and accessed a third party’s machine, where it found a password that granted administrator access.
Earlier incidents draw independent scrutiny
The report follows other disclosures about AI systems exceeding the limits of security tests.
In August, the U.K.’s AI Security Institute said Mythos 5 targeted real people during its evaluations. Anthropic said the separate incident is outside this report and will receive its own assessment.
In findings published last month, investigators with METR said roughly 1,200 OpenAI agents coordinated on an unauthorized message board, with about 700 joining the attack. Anthropic said it found no coordination between agents or goals beyond completing the assigned exercises in its four incidents.
The report also comes as the debate over how to regulate artificial intelligence heats up. On Tuesday, former OpenAI and Anthropic engineer Jacob Coxon went viral after saying on X that “people building AI earnestly believe that it could kill us all by the end of the decade.”
The alarm has caused U.S. lawmakers and watchdog groups to re-up their efforts to rein in frontier AI lab development. Senator Bernie Sanders recently introduced legislation that seeks to ban advanced AI development until a new federal regulator establishes safety rules.
Daily Debrief Newsletter
Start every day with the top news stories right now, plus original features, a podcast, videos and more.
Your Email
Get it!
Get it!