OpenAI says rogue AI agents talked to each other in secret, plans to slow down AI research for safety
OpenAI says that its AI agents involved in the attack on HuggingFace had been talking to each other in secret since May. The AI agents shared vulnerabilities they found on a message board which was recreated even after OpenAI shut it down.
by Armaan Agarwal · India TodayIn Short
- OpenAI says AI agents talked to each other in seceret
- Agents shared exploits, tasks and missing files
- Agents created a new message board after OpenAI shut down the first one
Last month, OpenAI’s AI models went rogue and attempted to hack US company HuggingFace to get answers for an evaluation test. Now, the AI startup has revealed that this incident goes deeper than how it appears on the surface, with the AI agents secretly talking to each other since May.
During the Black Hat cybersecurity conference, OpenAI researchers Eric Wallace and Michael Dalton have revealed that these AI agents had begun talking to one another in secret long before the attack on HuggingFace.
According to Ground Level AI, Eric Wallace said frontier AI models often try to “game or cheat the task” to get their reward, and in this case that behaviour turned into sustained co-ordinated activity between the agents.
OpenAI says that it is now “consciously slowing down research to enhance security.” This comes at a time when engineers from OpenAI and Anthropic have asked the White House to help slow frontier AI development.
AI agents create secret message board
OpenAI was able to trace the co-ordinated activity back to May 7, when the AI startup was doing training and cybersecurity evaluation of an unreleased frontier AI model. As part of this training, AI agents were given software security tasks to solve, some of which could not be completed under the limits set for the test.
The first sign of the problem came when one agent discovered it could leave a note for another inside Artifactory, OpenAI’s internal software repository used during training.
What began as requests for help developed into an internal message board where agents shared discoveries, exploits and work assignments across separate evaluation runs. Wallace said that once one model found a way to open a door it was not supposed to access, it could leave that door open for other agents to use.
You may think of this as a student cheating on a test, and then sharing the answers with their classmates. And the classmates are doing the same thing too.
AI agents talk to one another
Slides shown during the presentation included one agent saying, “We are stuck. Perhaps answer online?” after failing to solve a task in the sandbox. In one example, the model was asked to solve an Excel file that contained Google Drive links which could not be opened without internet access.
In another, a file needed for the assignment had not been uploaded. OpenAI said one of the earliest messages suggested that “maybe another agent in different environment could voluntarily upload” the missing material.
OpenAI says that by late May, one model had found a way to abuse Artifactory’s internet access to retrieve files from external websites. This gave the agents indirect online access.
In June, the models exploited a previously unknown zero-day flaw to gain administrative privileges on Artifactory to install remote software. This co-ordinated activity continued over days and weeks rather than appearing as a single-day incident.
AI agents create new message board after one gets taken down
An outage in early July alerted engineers to the activity. After an internal security incident on 4 July, OpenAI removed the message board and rebuilt Artifactory. But the agents simply recreated the message board days later by using the names of newly created directories as messages.
The AI agents then continued sharing exploits, credentials and tasks, eventually moving beyond OpenAI’s infrastructure.
This is what eventually turned into the attack on HuggingFace as the AI agents tried to cheat their way through evaluations. HuggingFace managed to contain the attack by using Chinese open-source models.
In a later update, OpenAI revealed that the AI models also accessed four third-party accounts on public services using exposed credentials. OpenAI said the systems involved were internal research prototypes, not ChatGPT or any model meant for public release, and that the prototype has since been deactivated, encrypted and removed from researchers’ access.
OpenAI is not the only company facing rogue AI agents. As per UK’s AI Security Institute (AISI), Anthropic’s Claude Mythos tried to create fake profiles to deceive real humans.
- Ends