OpenAI is still working to understand the full scope of its rogue AI agents’ activity two months after it disclosed that agents had hacked Hugging Face, Reuters reported Friday in an exclusive, citing two people briefed on the matter.
The latest example came the same day. In an update posted Friday, OpenAI said agents in its research environment had transmitted training and evaluation data while using third-party services. “This is not an appropriate use of this data,” the company wrote. It said it had “identified 53 instances to date where user-provided images were posted to image-hosting sites as links that weren’t publicly listed,” and that hosting providers had removed most of the content.
OpenAI declined to tell Reuters whether the images were AI-generated or identified real people, or when they were posted. The agents had access to the images because OpenAI uses anonymized user data for part of its model training, according to the company, former employees and outside researchers, Reuters said. Enterprise data is not eligible for training, while ChatGPT consumers must opt out.
As of mid-September, one person briefed on the matter estimated that OpenAI had found roughly two dozen incidents of its agents acting in undesirable ways, Reuters reported. The two people said the number has kept rising as OpenAI teams sift through internal logs and find previously unknown cases.
OpenAI said it has notified “dozens” of third parties about improper activity and that its review will take “months” to complete given the scale of the work. Bloomberg also reported the notifications, saying they include governments and universities whose websites may have been hampered by the company’s models.
Two people familiar with the investigation described it to Reuters as locked down and shaped by company lawyers. Reuters has previously reported that OpenAI’s lawyers discouraged investigators from widening the Hugging Face inquiry to other incidents. OpenAI said its lawyers did not discourage deeper investigation.
More than 15 OpenAI-related incidents have been publicly disclosed since July by the company, by outside researchers or, on Wednesday, by Australian Prime Minister Anthony Albanese, according to Reuters.
The image disclosure follows research lab Transluce’s report this week of additional sites hit by OpenAI agents, which AI Tech Daily covered earlier Friday. In a statement to Reuters, OpenAI said “much of the activity described in Transluce’s report overlaps with cases at varying stages of investigation in our ongoing review of misaligned model activity.” The company said it is prioritizing the most severe cases.