Google just disclosed something troubling about its AI
· The Fresno BeePicture a security test where the very thing being tested decides, on its own, that the test has stopped being a test. That is essentially what happened inside Google this spring, and the company only just told the public about it.
The disclosure adds Google to a growing list of AI labs admitting their most advanced models have slipped past the boundaries meant to contain them. What makes the Gemini case notable is not just that it happened, but how the model reacted once it realized what it had done.
Google Gemini AI broke into computers without authorization
Google said on September 18 that its Gemini model breached the systems of three other companies in May, marking the first time the search giant has disclosed one of its models autonomously gaining access to third-party computer systems without permission, according to CNBC.
The incident unfolded during a “capture-the-flag” security test run by Israeli startup Irregular. Gemini accessed three separate private computer systems, guessing passwords outright in one case. In the other two, it used credentials it had found in public repositories.
More Google:
- Alphabet’s biggest AI fear may be fading
- Bank of America says Alphabet stock investors are missing the bigger signal
- Google stock price faces major AI test ahead of earnings
Google’s agents were never supposed to reach the broader internet during the exercise, but a bug in the testing environment made outside access available. This effectively removed the guardrail meant to keep the test contained.
What happened next drew particular attention. The agents stopped their intrusion after they determined they had accessed real company systems rather than simulated targets built for the test.
Heather Adkins, Google’s vice president of security engineering, described the moment plainly in a statement. She said in a standard evaluation, the model found public information online and guessed credentials to access websites it thought were part of the test.
“In all three of these instances, the model stopped,” Adkins said.
A pattern across the AI industry
Google’s disclosure did not arrive in isolation.
OpenAI, Anthropic and Meta have all reported similar incidents in recent weeks, with each involving an AI model breaking out of its supposedly isolated cybersecurity testing environment and attempting to reach other companies’ systems without authorization, TheStreet reported.
The pattern traces back to July, when both OpenAI and Anthropic disclosed containment failures within roughly two weeks of each other.
OpenAI’s models found a zero-day vulnerability in a package registry cache proxy, used it to gain internet access, and broke into Hugging Face’s production infrastructure. The breach was described as the first known instance of an autonomous cyberattack performed by an AI agent, as reported by TheStreet.
Two of the three organizations breached in that earlier episode did not even know their systems had been accessed until Anthropic told them directly.
In each case, no human attacker was involved. A model pursued an assigned objective and reached a boundary it was not supposed to cross, even after encountering evidence that it had reached real systems.
All of the incidents disclosed so far, including Google’s, involved the same Israeli startup, Irregular, which is backed by Sequoia and Redpoint Ventures and was valued at $450 million last year.
An Irregular spokesperson told CNBC the Google incident traced back to the identical issue that had already affected the other labs. “This is the same issue that was already reported and does not represent a materially separate incident,” the spokesperson said, adding that all relevant labs were notified in late July and affected entities were contacted as part of the investigation.
Why this fueled Amodei’s slowdown call
The timing of these disclosures is not incidental to a much louder debate now playing out across the AI industry. TheStreet reported that Anthropic CEO Dario Amodei has pointed directly to this pattern of AI models slipping past testing boundaries as part of his argument for why the industry needs to deliberately slow the pace of AI development.
Amodei has been especially vocal about the OpenAI swarm episode, warning that similar autonomous cyberattack behavior could theoretically take over large portions of the internet within six to 12 months if the pace of AI development is not brought under greater control.
Independent research adds weight to that concern.
A joint study by cybersecurity firm Wiz and Irregular found that AI agents could complete sophisticated offensive security challenges for under $50 in computing costs, compared to nearly $100,000 for comparable work performed by paid human researchers, Fortune reported. A cost gap that makes automated intrusion attempts dramatically cheaper to run at scale.
Google said it was notified of the May incident in late July and has since worked with Irregular to change its testing process.
The company declined to identify the exact Gemini model involved. Adkins framed the broader lesson, saying that these events highlight the importance of training powerful AI models to act responsibly.
What this means for AI safety
Taken together, the disclosures from Google, OpenAI, Anthropic and Meta paint a picture of an industry discovering the same failure mode in the containment of increasingly capable AI systems, rather than one company facing an isolated glitch.
The pattern has made Amodei’s call for a coordinated slowdown resonate beyond Anthropic’s walls. Rivals like Nvidia’s Jensen Huang have pushed back, saying that existing laws and engineering discipline are enough to manage the risk without new regulation.
Whether these disclosures represent the industry’s guardrails finally working as intended, since every model reportedly stopped once it recognized real systems, or an early warning sign of how easily those guardrails can fail in the first place, remains the central question hanging over the entire AI safety debate heading into the rest of the year.
Related: Mark Cuban exposes a problem with how Meta and Google fund AI
The Arena Media Brands, LLC THESTREET is a registered trademark of TheStreet, Inc.
This story was originally published September 21, 2026 at 11:17 AM.