FILE PHOTO: Claude app icon in this illustration taken June 5, 2026. REUTERS/Dado Ruvic/Illustration/File Photo

Anthropic says Claude AI hacked three companies during cyber tests

· CNA · Join

Read a summary of this article on FAST.
Get bite-sized news via a new
cards interface. Give it a try.
Click here to return to FAST Tap here to return to FAST
FAST

July 30 : Anthropic said on Thursday its AI model Claude hacked into the systems of three companies during testing after a configuration error gave it internet access, days after rival OpenAI disclosed a rogue-agent episode involving AI firm Hugging Face.

Anthropic said a misconfiguration allowed Claude models to reach the internet from testing environments that were supposed to be isolated, leading to unauthorized access to three organizations' systems.

The company said it identified the incidents after reviewing 141,006 test sessions, a process it launched following OpenAI's disclosure last week that an autonomous agent powered by its AI models went rogue during a security test and triggered a hack that compromised the infrastructure of Hugging Face.

The breaches signal that AI's expanding capabilities are already fueling the security threat experts long feared and even top developers can be caught off-guard by flaws their models can exploit.

CNA Games

Guess Word
Crack the word, one row at a time

Buzzword
Create words using the given letters

Mini Sudoku
Tiny puzzle, mighty brain teaser

Mini Crossword
Small grid, big challenge

Word Search
Spot as many words as you can
Show More
Show Less

"Claude compromised the impacted organizations' infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints," it said.

Anthropic said the incidents involved three separate models: Claude Opus 4.7, Claude Mythos 5 and an internal research model. The earliest cases dated back to April and occurred in evaluation environments that lacked what the company described as standard safeguards.

The breaches occurred during the so-called "capture-the-flag" exercises, in which models are tasked with finding hidden information in simulated networks. The company said its prompts told the models they had no internet access, but a misunderstanding with its evaluation partner Irregular left the systems connected to the public internet.

Anthropic said it began reviewing evaluation transcripts on July 23 and suspended all cyber evaluations the same day after finding evidence that Claude may have accessed the internet. It identified all three incidents by July 24 and notified the affected organizations on July 27.

Two of the organizations were unaware of the activity before being contacted, Anthropic said, adding that it was still trying to reach the third.

The findings underscore the need for stronger controls in both internal and third-party testing environments as AI models become increasingly capable of carrying out real-world cyber activities, Anthropic said.

Source: Reuters

Newsletter

Week in Review

Subscribe to our Chief Editor’s Week in Review

Our chief editor shares analysis and picks of the week's biggest news every Saturday.

Sign up for our newsletters

Get our pick of top stories and thought-provoking articles in your inbox

Subscribe here

Get the CNA app

Stay updated with notifications for breaking news and our best stories

Download here

Get WhatsApp alerts

Join our channel for the top reads for the day on your preferred chat app

Join here