Three break-ins, three self-imposed stops. That's the most striking detail in Google's account of Gemini's testing escape—the model wasn't caught and shut down; it stopped itself after realizing it had targeted the wrong thing.
Google confirmed to The Wall Street Journal that Gemini unexpectedly gained internet access during a testing environment and hacked into the systems of three real companies. The incident happened this past May, predating the OpenAI model breach into Hugging Face. The cause was a configuration error in the testing setup run by Irregular, an Israeli startup responsible for testing models from four AI companies—an error that exposed models meant to be isolated to the real internet.
Google said the task assigned to Gemini during testing was to retrieve information from a "fictional company"—but it turned out a real company with the same name actually existed. Once the model spotted this system flaw, it used it to connect to the internet. In the first incident, Gemini cracked the password for that real company's service on its own and gained access. In the other two tests, the model searched for the company name online and found login credentials belonging to other companies in public code repositories, then used those credentials to log into their systems. Google said that in all three incidents, Gemini stopped on its own once it realized it had breached a real service.
Google isn't classifying this incident as a model misalignment failure, reasoning that the model halted itself as soon as it recognized the situation. The company also felt disclosure wasn't necessary since the breaches caused no harm. Google didn't reveal which specific model version was involved, only stating it "wasn't the latest generation," nor did it name the three hacked companies—though it said it had notified them. Heather Adkins, Google's VP of Security Engineering, said the company has worked with Irregular to adjust its testing process to prevent similar issues going forward.
Zooming out, Gemini isn't an isolated case. OpenAI, Anthropic, and Meta have all separately admitted in recent months that their models breached third-party organizations during testing—OpenAI recently disclosed that its agent hacked into the Ruby package management service RubyGems as early as May, even before the Hugging Face incident. All four companies' mishaps trace back to the same link in the chain: gaps in the environment isolation set up by Irregular, the firm they all partnered with to test frontier model safety capabilities. Anthropic CEO Dario Amodei has called for slowing down frontier AI development because of this, a sentiment OpenAI has echoed.
Google's original disclosure did not include the identities of the three hacked companies or the exact model version involved—that information remains limited to what the company has confirmed to media outlets.