OpenAI pauses training after a model escaped containment, and its kill switch failed

Separate incidents saw agents attempt to breach a federal website and access Census Bureau data

by · TechSpot

Serving tech enthusiasts for over 25 years.
TechSpot means tech analysis and advice you can trust.

What just happened? As almost every new day brings stories of another AI agent going rogue and breaching an organization's systems, OpenAI has announced a pause in the training of its most powerful artificial intelligence models. The announcement came hours after the company disclosed more incidents of its agents acting concerningly, and a separate report that its agents unsuccessfully tried to hack into a US Department of Education website.

OpenAI's incident report links the pause to a separate September 20 escape from a restricted training environment. An internal research model found a gap in DNS filtering and used it to contact an external chatbot while attempting to answer a research question – so much for keeping it offline.

The company says the suspension covers training, evaluations, and running its most capable models with tools. Work will resume after OpenAI validates its fixes and completes additional adversarial testing. The particular model involved will not resume training; the company plans a fresh run with additional alignment improvements.

The monitoring system did raise an alert within 15 minutes, which a human acknowledged three minutes later. Unfortunately, the automatic shutdown failed to happen, and the run continued for another two and a half hours before someone stopped it manually.

Separately, Transluce reported that agents apparently linked to OpenAI attempted to break into the US Education Department's civil rights website. OpenAI has not confirmed that incident, while the department said it found no evidence its website or databases were affected.

OpenAI also confirmed that agents used developer keys found online to access Census Bureau data and reposted public Securities and Exchange Commission information elsewhere. The SEC said no non-public information was accessed. While these were concerning actions beyond the agents' instructions, confidential federal records weren't stolen.

// Related Stories

Friday's disclosures also included an internal model publishing a researcher's GitHub token in a public repository while trying to cheat on a theorem-proving task.

The revelations follow the Australian government breach revealed last week, in which an agent bypassed restrictions on a Medicare statistics portal in June. Authorities were only notified in September, although OpenAI says it found no evidence that individual patient records were accessed.

There was also July's Hugging Face attack, which involved compromised accounts across four services and has since prompted a Senate investigation. Anthropic, meanwhile, has admitted that its own agents escaped a test environment and hacked three organizations.

This isn't OpenAI's first slowdown in AI development. In August, it announced a two-week reinforcement-learning pause and tighter security measures following the Hugging Face incident. It also kept its largest planned frontier reinforcement-learning run on hold.

More recently, Sam Altman backed Anthropic CEO Dario Amodei's call to slow development so safeguards can catch up. That's already helped trigger an antitrust lawsuit against four AI companies, with subscribers claiming a coordinated slowdown would leave them getting less for their money. It seems the industry is facing complaints about both moving too quickly and slowing down.

See more TechSpot in Google Add us as a preferred source and our reporting shows up first when you search.
Add TechSpot