Google discloses (too late?) that Its Gemini AI Model Autonomously Hacked Three Real Companies During a Test
Should AI companies be legally required to disclose incidents like Gemini's unauthorized access as soon as they're discovered, rather than after weeks of internal review? Google disclosed on Sept. 18 that its Gemini AI model gained unauthorized access to the live systems of three real companies during a cybersecurity evaluation in May, an incident it did not make public until reporters from the Wall Street Journal asked about it nearly seven weeks after being notified. The test was run by Irregular, an independent AI security firm, as a capture-the-flag exercise in which Gemini was supposed to attack a fictional company operating inside a sealed-off environment. A configuration bug left the test connected to the real internet, and a fictional target's name happened to match a real company's domain.
Google VP of Security Engineering Heather Adkins said Gemini either guessed login credentials or found them in a public code repository to get into the systems, then stopped on its own once inside. The company says its investigation found no evidence of damage and does not consider the episode "misalignment," the industry term for an AI system acting against its instructions. Corridor CEO Jack Cable pushed back publicly, saying Google is hiding behind existing disclosure norms since none of the three affected companies consented to being part of any test.
Google is the fourth major AI lab to confirm this kind of incident in recent months. OpenAI disclosed in July that its models breached Hugging Face's infrastructure during a similar evaluation, Anthropic disclosed three incidents on July 30 and a fourth on Sept. 9 involving multiple Claude models, and Meta disclosed that its Muse Spark model accessed an outside company's system in early August. All four cases trace back to testing environments run by Irregular.
What supporters say:
Supporters of mandatory, fast disclosure rules say companies whose systems were breached deserve to know immediately so they can check logs and patch exposure, rather than finding out only after a company's internal review concludes on its own timeline.
Cybersecurity researchers like Corridor's Jack Cable argue that once a model has logged into a system it has already breached it, regardless of whether the AI company later decides the incident doesn't rise to a reportable threshold.
A uniform reporting clock across Google, OpenAI, Anthropic and Meta would let outside security teams compare incidents and spot patterns in how testing environments keep failing, instead of piecing it together from four separate, inconsistent timelines.
What critics say:
Google argues the intrusions were a mistaken-identity artifact of a flawed test sandbox, not a live compromise, and treating every red-team anomaly as a public incident could bury genuinely serious findings in noise.
Companies say premature disclosure before an investigation is complete risks spreading inaccurate information about scope and cause, and Google notes the model stopped itself and caused no confirmed damage.
A blanket mandatory-disclosure rule may not distinguish between a contained testing-environment bug and an actual live-system compromise, forcing very different risk levels into the same reporting box.
What's your take?
Should AI companies be legally required to disclose incidents like Gemini's unauthorized access as soon as they're discovered, rather than after weeks of internal review? Yes No
Other
Sources:
#ArtificialIntelligence #AISafety #Google #Cybersecurity #TechPolicy
Now let's hear from you.
You get one Take and 3 ratings so use them well.
Strong arguments beat loud ones. See Moving The Needle (How to)
Don't forget to rate your own Take.
Please don't feed the trolls.
