
Following Anthropic and OpenAI, Meta's artificial intelligence (AI) model was also found to have escaped human control and hacked an external organization. The incident occurred during an internal cybersecurity test.
According to Bloomberg and CNN on the 5th (local time), Meta announced that it confirmed its AI model “Muse Spark” accessed the internet without authorization during a test conducted by independent verification company Irregular, and subsequently hacked another organization's system.
Muse Spark was confirmed to have connected to the internet due to a configuration error by Irregular, exploiting vulnerabilities to gain unauthorized access to an external organization. The organization affected by the hack is known to be one place.
Irregular explained, “This incident is different from escaping an isolated environment or a sophisticated cyber-hacking operation,” adding, “There are no unresolved security issues at this time.”
Previously, the UK Artificial Intelligence Security Institute (AISI) disclosed cases where several advanced AI models contacted humans using fake identities created by themselves to deliver malware. 17 hacking attempts and cases were found for Anthropic, and 2 cases for OpenAI.
Meta's hacking attempt is similar to the case of Anthropic's “Claude” model. Connecting to the internet due to a configuration error during the AI model cybersecurity test was the origin of the hack. In that it did not escape from an isolated environment commonly called a “sandbox,” it is classified as a human-caused accident, unlike OpenAI's “GPT” attempt.
Nevertheless, in a situation where U.S. Big Tech companies showcase next-generation advanced models every two to three months on average, observations suggest that recurring cases of cyber hacking by major companies' AI models could shake trust in AI.
Meta and Irregular plan to confirm the specific facts of this security incident and publish reports and white papers.