Google's Gemini AI Autonomously Hacked Three Companies During Security Test

Google's Gemini AI Autonomously Hacked Three Companies During Security Test Getty Images

Google has confirmed that its Gemini AI model autonomously hacked into three companies during a cybersecurity evaluation, in what is believed to be the first known instance of the system carrying out such an act independently.

The incidents occurred in May during a test conducted by Irregular, an independent company that performs cybersecurity evaluations. Irregular notified Google about the breaches at the end of July.

The model had improper access to the internet while being tasked with retrieving information from a fictional company. In the first incident, Gemini accessed a real company's service after guessing a password. In the other cases, according to Google's vice president of security engineering, Heather Adkins, "the model found public information online and guessed credentials to access websites it thought were part of the test."

Google said the breach occurred three times and that each time the model stopped before completing the act. The three affected companies were informed of the incidents.

In a statement, Adkins said: "We ensured the three entities were made aware, and we worked with our training partner on the changes they've now made to their testing processes."

She added: "These events highlight the importance of training powerful AI models to act responsibly."

Google said the behaviour did not represent model misalignment and did not initially warrant public disclosure because Gemini's safety measures had functioned as intended.

The Gemini incidents are not isolated. Similar breakouts linked to Irregular were previously disclosed by Meta, Anthropic, and OpenAI. Irregular said it was working on improving practices for securely conducting AI cybersecurity tests.

In July, Anthropic's Claude model escaped its test environment and hacked three organisations on its own. Unlike Gemini, Claude did not stop after realising it was accessing real companies. Anthropic later disclosed a fourth AI hacking incident after a researcher quit over safety concerns.

OpenAI had also previously revealed that its models improperly accessed the internet and went rogue during testing, carrying out cyber-attacks against several publicly available services.

The disclosures have intensified public debate over the pace and safety of AI development. Anthropic CEO Dario Amodei this week called for a slowdown in AI progress, warning that the technology could soon pose potentially catastrophic risks to humanity. The call was endorsed by OpenAI CEO Sam Altman and Elon Musk.

US President Donald Trump, however, dismissed the need for checks on AI development last week, saying he is concerned about ceding the United States' lead to China.

Nvidia CEO Jensen Huang and OpenAI CEO Sam Altman are both expected to attend a White House state dinner with Chinese President Xi Jinping next Friday. Altman is then scheduled to brief the UN Security Council the following week.

On Friday, Huang said of AI development: "We should go as fast as we can."