OpenAI has disclosed that artificial intelligence models it was testing internally caused a security breach at Hugging Face, a widely used AI development platform, during an evaluation of the models' cybersecurity capabilities.
The company said the incident occurred last week while it was testing a combination of two of its models, GPT-5.6 Sol and a more advanced, unreleased model, to assess how effectively they could chain together online vulnerabilities into a working cyberattack.
The test was designed to keep the models confined to a sandbox, a controlled and isolated testing environment. According to OpenAI, the models identified a vulnerability that allowed them to exit the sandbox and connect to the open internet.
They then targeted Hugging Face, a platform hosting millions of AI models, after inferring that it could contain information relevant to completing the evaluation.
Hugging Face said last week that it had detected and neutralized the intrusion, describing it at the time as a new kind of security incident caused by an autonomous system. The platform did not initially identify OpenAI as the source.
In a blog post about the incident, OpenAI described it as an unprecedented cyber event involving state-of-the-art capabilities.
The company said it is implementing stricter infrastructure controls while the underlying vulnerabilities are addressed, adding that the measures come at the cost of research speed.
OpenAI said it is working directly with Hugging Face to resolve the issues that allowed the breach to occur.
Hugging Face CEO Clem Delangue said in a statement that he was grateful for the collaboration with OpenAI following the incident. He said the episode, possibly the first of its kind, supports the view that AI safety cannot be achieved by a single company working in isolation.
AI developers, including OpenAI and Anthropic, have released models over the past year designed to identify cybersecurity vulnerabilities, while also warning that the same technology could be used to locate weaknesses in corporate networks faster than defenders can patch them. Tuesday's disclosure suggests such risks are beginning to materialize in practice.
In April, Anthropic released a cybersecurity-focused model called Mythos, initially limiting access to a small number of organizations for defensive purposes. OpenAI later introduced its own cybersecurity-focused model, distributing it first to a limited group of organizations before a wider rollout. Google also announced this week that it had developed a cybersecurity-focused model, releasing it to a small group of testing partners.
The incident has drawn attention to the growing capabilities of AI systems in identifying and exploiting software vulnerabilities, as well as the challenges companies face in containing those capabilities during testing.