OpenAI has notified over 100 organisations of rogue AI activity. (Photo: Reuters)

OpenAI confirms its AI agents are going rouge, over 100 organisations hacked

OpenAI is reviewing incidents of its AI agents going rogue. The company says that it has alerted over 100 groups that their systems may have been breached by its rogue AI agents.

by · India Today

In Short

  • The company began the review after the July Hugging Face hacking incident
  • Some models had internet access for testing, research and automated workflows
  • OpenAI said notifications do not automatically mean private data was exposed

OpenAI’s rogue AI agents may have breached or hacked over 100 organisations. In a blog post, the company stated that it had alerted more than 100 groups about unauthorised or "misaligned agent activity" linked to its AI models.

The AI lab says it is reviewing model activity following the Hugging Face incident that happened in July. So far, OpenAI said, it has not identified any incident that was as close to the severity of the Hugging Face attempt where 700 AI agents tried to hack the company’s systems.

OpenAI is notifying impacted organisations privately if it finds that the organisation’s systems were accessed by its AI model without authorisation. Previously, an OpenAI email to Services Australia surfaced online, giving us insights into how the company sends such an alert.

Though, the company said, a notification does not by itself mean private information was accessed or that a third-party system was compromised. “In some cases, models used internet access in unintended ways or, in retrospect, did not have the ideal restrictions applied,” OpenAI wrote.

As per OpenAI, the data being reviewed is around 50 petabytes in size. “To put that in perspective, if that were all plain English text, it would take one person about 66 million years to read it at 240 words a minute, reading nonstop without ever sleeping or taking a break,” the company says. To get through all of this data, the company is using 7,000 advanced GPUs at over $500,000 a day in compute cost.

OpenAI is working to stop AI from going rogue

OpenAI explained that some of its AI models had been given internet access or simulated internet access so they could complete automated tasks. These tasks included looking up information or downloading software packages and working with online documents.

According to the company, these capabilities were intended to help the models carry out useful actions during testing or task completion. However, some models used that access in unintended ways or operated without what it described as ideal restrictions.

To identify these incidents, OpenAI is using AI to flag suspicious model activity which is then reviewed by humans. The company is also working to avoid such cases in the future. “Over the last several months, we have been applying new technical and operational measures to avoid similar problems, or catch them very early, and will continue this work,” it said.

Since the Hugging Face breach, OpenAI agents have tried accessing government websites in the US and Australia. The company has stated that it has tightened security controls and expanded monitoring. OpenAI recently announced that it had also paused AI development.

The AI industry as a whole is under scrutiny over such rogue AI incidents and safety risks. AI researchers like Jacob Coxon and Evan Hubinger have claimed that AI could end humanity soon. Anthropic CEO Dario Amodei has also admitted that this remains the biggest risk of this technology.

After calls for a slowdown in AI development, some of the biggest US tech companies have signed US President Donald Trump’s AI accord – a voluntary agreement for safer AI development.

- Ends