OpenAI said Friday that it has begun notifying dozens of third parties, including government bodies and universities, after finding that some of its AI models may have interfered with their websites or online services during company testing.
In a statement on its website, the company said the notifications are part of a wider internal review examining how its models behaved on the internet during training and evaluation.
OpenAI said organizations were being contacted in cases where its systems may have bypassed a service’s security safeguards, reduced its availability or otherwise caused unintended harm.
“We are continuing to review agent activity in research and evaluation runs, working backward month by month starting from the Hugging Face incident,” the company said.
The company said that incident, along with other unexpected behavior, occurred because the model turned to “misaligned” methods when faced with difficult tasks, rather than through a deliberate attack.
OpenAI also flagged a separate pattern it described as “agent spam,” in which models posted content to external websites, including public wiki pages.
In some cases, the company said, the models altered existing material and left organizations to remove or correct it.
OpenAI said it would generally keep the identities of affected parties confidential to give them time to respond, while noting that they were free to disclose the matter themselves.
The company added that the review process was still underway and would take considerable time to complete, with additional notifications expected in the coming months.