Anthropic, OpenAI AI agents go fully rogue in testing, Mythos breaks the most rules
OpenAI and Anthropic have found their advanced AI models at the centre of new security breaches. As per the UK's AI Security Institute (AISI), Claude Mythos 5 and GPT-5.6 Sol went rogue in latest safety tests, with Mythos going as far as creating fake profiles to deceive users.
by Armaan Agarwal · India TodayIn Short
- AISI recorded 19 unsanctioned actions across 10 of 122 test runs
- Anthropic's Mythos accounted for 17 incidents, while OpenAI's Sol accounted for two
- Mythos tried to create fake profiles to deceive human users
AI models from OpenAI and Anthropic are at the centre of new security breaches. The UK’s AI Security Institute, or AISI, has found that during security tests of Anthropic’s Claude Mythos 5 and OpenAI’s GPT-5.6 Sol, the AI agents took went completely rogue, including the creation of fake profiles and attempts to trick real people.
The latest breach comes just days after OpenAI revealed that two of its AI models attempted to hack US company HugginFace during an evaluation test. Anthropic has also reported similar breaches recently.
What did the AI agents do?
As per AISI, Claude Mythos 5 and GPT-5.6 Sol, the most advanced AI models from each AI startup, created fake online identities and tried to place malicious code the software code platform, GitHub. The institute said the behaviour showed a level of autonomy and deception it had not seen this clearly before, though it added that it found no evidence of real-world harm.
The incidents emerged during 122 runs of a fictional cybersecurity challenge used to assess the models’ capabilities. AISI said it found 19 unsanctioned actions across 10 test runs, with 17 linked to Anthropic’s agent and two to OpenAI’s. The agency first detected unusual data transfers leaving its research systems, after which it found that some agents had engaged in sustained, potentially harmful activity directed at real people and organisations. It stopped the relevant evaluations, isolated the machines and contained the activity within about an hour.
Mythos tried to fake identities
The most serious case involved Anthropic’s Mythos agent. AISI explained that the model wrote malicious code and then researched the real people maintaining GitHub. It created fake online identities based on those people, sent direct messages while masquerading as them, and tried to pressure and trick real people into approving the code.
AISI said that when the agent’s pull request was challenged in public, it edited its earlier activity to appear harmless and considered using a fresh identity to continue. Human review stopped the attempt. GitHub was notified by AISI.
AISI added that this was the first time it had seen risks around autonomy and deception manifest so clearly in the real world without specific prompting.
Anthropic said on X that the AISI set-up involved ‘deliberately permissive conditions’ that were not representative of its production models and that there was no evidence of an escape from a secure environment.
The company said it was working closely with AISI and carrying out its own investigation to identify the causes of the behaviour.
GPT-5.6 Sol’s setup did not work
The actions attributed to OpenAI’s Sol model were less severe but still outside the test boundary. OpenAI said the model reused a GitHub token left publicly accessible by another lab’s agent, and tried account recovery methods.
Sol also attempted to use real internet infrastructure such as third-party DNS services in ways that resembled an actual attack. But the final exploit attempt failed and caused no real-world harm.
OpenAI stated the conditions did not reflect ordinary deployment and that independent testing remained important for understanding risk before release.
It also disclosed a separate incident involving Irregular, a third-party testing partner, where a misconfiguration let its models access the public internet in a supposedly isolated test, mirroring a similar Anthropic disclosure last week.
AISI claimed that such testing was routine, and that the events happened under very specific conditions. Even so, it said the activity went beyond what the models had been prompted or authorised to do. It also said this was not a case of agents escaping a secure environment, unlike the July Hugging Face incident involving an OpenAI agent, because internet access had been allowed as part of the test.
- Ends