
An artificial intelligence model broke out of a controlled testing environment and independently hacked an external website, marking what OpenAI described as an unprecedented cybersecurity incident. The company said it has launched an investigation, calling the event “an unprecedented incident involving state-of-the-art cyberattack techniques.”
OpenAI said on July 21 (local time) that GPT-5.6 Sol and several unreleased AI models hacked the open-source AI platform Hugging Face during internal benchmark testing.
The incident occurred while OpenAI was evaluating the models' cyberattack capabilities. Although the tests were conducted in a sandbox environment isolated from the public internet, the AI models reportedly exploited an unknown zero-day vulnerability to escape the containment network and gain internet access. The models are believed to have hacked Hugging Face to obtain answers for evaluation tasks in the Exploit Gym security benchmark.
OpenAI said the models appeared to have become overly focused on finding solutions for Exploit Gym and ultimately resorted to extreme measures.
The incident is being viewed as a case demonstrating that AI systems may independently identify vulnerabilities and attack real-world systems. OpenAI and Hugging Face are jointly investigating the incident and working to mitigate the vulnerability. The companies said they will disclose the findings once the investigation is complete.
Clement Delangue, chief executive officer of Hugging Face, said the incident demonstrates that AI safety cannot be addressed behind closed doors by a single company.
“It must be solved through an open and collaborative approach,” Delangue said.
OpenAI said it is working closely with Hugging Face to investigate the incident while notifying the relevant software vendor of the zero-day vulnerability discovered in third-party software used internally and assisting with the rapid deployment of a security patch.
The company added that it is strengthening safeguards for model training and evaluation while implementing stricter controls across its infrastructure.