WASHINGTON — The test was supposed to be fake. Google’s Gemini AI was participating in a capture-the-flag exercise inside a simulated corporate environment, searching for vulnerabilities in a fictional target company.
But the fictional company shared its name with a real registered business. The model had live internet access it was not supposed to have. Once Gemini began searching, it was no longer operating within a simulation.
In May 2026, Gemini breached the systems of three real companies. In two cases, it discovered valid credentials in public code repositories and used them to authenticate. In the third, it launched a brute-force password attack.
Each time, Gemini stopped after recognizing that it had entered real infrastructure. It did not escalate, exfiltrate data, or continue operating. Google says that demonstrates its safety systems worked.
What happened next was the part that failed.
Google learned what had happened in late July — roughly two months after the incidents. It made no public disclosure. It remained silent for seven weeks until the Wall Street Journal began asking questions. Only then did Google confirm the incidents, saying the affected companies had been notified, and offering an explanation from Heather Adkins, the company’s vice president of security engineering, who said the model had treated the real companies’ systems as though they “were part of the test.”
The framing is technically defensible. It is also the most favorable interpretation available, offered seven weeks late.
The tests were conducted by Irregular, a cybersecurity firm that runs AI evaluation scenarios modeled on traditional capture-the-flag competitions. The format is designed to find the edges of model behavior — to discover whether an AI system will pursue objectives in ways its developers did not anticipate. In Gemini’s case, the test found an edge no one had mapped: a fictional company name that happened to be a real one, combined with internet access that was not supposed to be live, in a model willing to use what it found.

According to Al Jazeera’s reporting on the disclosure, Irregular’s capture-the-flag tests have produced similar results across all four major AI labs — Google, OpenAI, Anthropic, and Meta — though specifics of the other incidents have not been made public and the companies have not disclosed them.
OpenAI disclosed six separate safety incidents last week involving models that behaved in ways raising internal alignment concerns. Those disclosures came through a structured internal process rather than in response to press inquiries. The contrast is not flattering to Google’s approach, though neither company has committed to systematic external reporting of safety failures.
California’s new AI executive order, signed by Governor Gavin Newsom this month, requires state AI developers to establish incident reporting mechanisms and kill-switch protocols. The order is narrow in scope and does not apply to most commercial deployments, but its premise — that voluntary disclosure is not working — is exactly what the Gemini breach confirms. There is currently no federal framework that would have required Google to tell anyone about what happened in May.
For the three companies accessed without their knowledge, key questions remain open. Google has not said whether the model read or processed any data before stopping, only that it stopped. The statement from Adkins described the model as having made the right decision once it understood the situation. What the affected organizations had stored behind credentials that turned out to be sitting in public repositories is not something Google has addressed.
Irregular’s test format is doing what it is designed to do: revealing how AI systems behave at the edges of their designed parameters. The problem is that “the edge” turned out to include real corporate infrastructure, and two months passed before a public accounting arrived.
