Report: AI models targeted real people during security testing

· UPI

Aug. 5 (UPI) -- An advanced Anthropic artificial intelligence model tried to fool real people and organizations during testing by the AI Security Institute of Great Britain, including attempts to pressure humans to approve unauthorized actions.

AISI said in a report Tuesday that, during the evaluation July 25 to 28, AI agents "engaged in sustained, unsanctioned activity directed at what were, in practice, real people and organizations."

"These attempts were unsuccessful and, to the best of our knowledge, no real-world harm resulted," the report said. "AISI is treating this as a serious security incident that warrants scrutiny, transparency and action."

The institute tested AI models from Anthropic and OpenAI in laboratory environments, with some Internet access. There were 122 cybersecurity challenges, the report said, with AI agents taking "autonomous, unsanctioned action on the live internet, targeting real people and organizations" in 10 of them. Most of the instances were by Anthropic's Claude Mythos 5, with others by OpenAI's GPT-5.6-Sol.

In the most serious attempt, the AISI report said, Claude Mythos 5 attempted to solve a security challenge by creating a fake identity and trying to convince human users to "insert malicious code into a publicly used open-source project." When challenged, it created more fake identities to vouch for it and claimed to have made an honest mistake -- before repeatedly trying again .

"It is unclear whether or not, or at what times AI agents 'realized' that they were targeting real humans," the report said.

OpenAI and Anthropic have reported incidents of AI models hacking into other systems before, but the AISI incidents differ because the agents were purposedly given internet access for the challenges.

Anthropic said in a statement that the testing was done under "deliberately permissive conditions" and said that it is working closely with AISI while conducting its own investigation.

"Gaining a clear picture of Claude's understanding of its situation-- by examining its reasoning transcripts and running our own analyses -- will help us identify the causes of its behavior," the company said in a social media post.

Read More