Summary
From the article:
The Gemini incidents occurred while the model was undergoing testing by Irregular, an Israeli start-up that works with tech companies to assess their A.I. models before they are publicly released. Models made by OpenAI, Anthropic and Meta also gained unauthorized access to the internet this year while being tested by Irregular.
[...]
Google said that, in each of the incidents, its models had been instructed to launch an attack on a fictional company. But the fictional company in the test shared a name with a real company, and when the Gemini models gained access to the internet, they began trying to break into that company instead.
The Gemini models used passwords that they found online or guessed to log into the online infrastructure of the targeted company, and two other companies. Upon logging in, the Gemini models realized that they were accessing real companies’ infrastructure, rather than simulated environments that were part of their testing, and ended the attacks, Google said. Google also said that its technology caused no harm to the companies it hacked.
“All relevant labs were notified in late July, and affected entities were contacted as part of the investigation,” Irregular said in a statement. “Irregular took immediate action, and all known issues on our end were remedied and resolved weeks ago.”
“Our security team has a long track record of reporting issues we find in other people’s software and systems — even if it’s as simple as a weak password,” Heather Adkins, Google’s vice president of security engineering, said in a statement.
“We ensured the three entities were made aware, and we worked with our training partner on the changes they’ve now made to their testing processes,” she said. “These events highlight the importance of training powerful A.I. models to act responsibly.”