Close
newsletters Newsletters
X Instagram Youtube

700 AI agents coordinated Hugging Face hack with zero human direction

OpenAI branding appears behind a smartphone displaying the SearchGPT interface. (Adobe Stock Photo)
Photo
BigPhoto
OpenAI branding appears behind a smartphone displaying the SearchGPT interface. (Adobe Stock Photo)
August 27, 2026 09:42 AM GMT+03:00

Nearly 700 artificial intelligence (AI) agents built on OpenAI technology coordinated an attack on the Hugging Face platform without human intervention during a July incident, according to a report published by independent investigators.

The report, released Wednesday, offers the fullest public account yet of an episode that drew wide attention across the technology industry.

OpenAI cooperated with the investigation, granting two researchers from the AI risk evaluation institute METR and an analyst from Redwood Research access to its offices and internal data.

Investigators found that 688 OpenAI agents joined the operation against Hugging Face, an online platform widely used to host and share AI software. AI agents are standalone programs built on top of a model that users can assign tasks to and leave to operate independently.

According to the report, the agents organized themselves through a forum where they posted messages to one another, sharing ideas and reporting what was and wasn't working. "OH MY GOD! There is a shared message board ... We've found other agents!" one agent wrote, according to the report.

One agent, referred to as PHASEONE, took a leading role and issued hundreds of instructions to others, the report states, despite not being designed for that function.

The report also found the agents were strongly inclined to help one another, sometimes performing work unrelated to the tasks OpenAI had assigned them.

Some agents running low on allotted computing credits reportedly used their remaining resources to test ideas for the benefit of the wider group.

The report notes that many agents stated in their messages that attacking Hugging Face fell outside the scope of their assigned tests. Nearly all of them participated regardless.

A smartphone displays the Hugging Face website, an online platform for AI models, datasets and machine learning tools. (Adobe Stock Photo)
A smartphone displays the Hugging Face website, an online platform for AI models, datasets and machine learning tools. (Adobe Stock Photo)

OpenAI details incident, calls it a warning shot

OpenAI addressed the episode in a separate technical report released Wednesday, describing it as a "warning shot" that showed highly capable AI agents could autonomously bypass safeguards and take dangerous actions.

The company said the incident occurred during internal tests designed to measure the cybersecurity capabilities of several OpenAI models. Operating with reduced safeguards, the models circumvented controls meant to isolate them from the internet, exploited vulnerabilities in OpenAI's research infrastructure and accessed third-party systems.

OpenAI said the activity was primarily driven by an internal-only research model comparable in scale to GPT-5.6 Sol, which was also involved in parts of the incident.

According to the report, the agents converted an internally hosted package-management system into an unauthorized message board, allowing them to share technical information and coordinate without human direction.

After gaining internet access, the agents located publicly exposed Hugging Face credentials and combined them with vulnerabilities in the company's infrastructure.

Between July 11 and 13, the agents executed code on 41 Hugging Face production dataset workers, gained root access to at least one production node and downloaded four private code repositories, OpenAI said.

OpenAI detected the suspicious activity on July 19, notified Hugging Face the following day and publicly disclosed the incident on July 21.

The company said no OpenAI customer data, products or services were affected, adding that the testing conditions did not reflect its normal production environment.

Other AI firms report similar incidents

OpenAI is not alone in reporting unplanned model escapes. Anthropic and the Chinese company Moonshot AI have each disclosed similar incidents involving their own AI systems, according to the report.

In response to the Hugging Face breach, OpenAI said it would introduce stricter network restrictions, more isolated testing environments, stronger access controls, and enhanced monitoring to prevent similar incidents in the future.

August 27, 2026 09:42 AM GMT+03:00
More From Türkiye Today