OpenAI is conducting a broad review of AI agent activity after the Hugging Face incident, examining whether models exceeded assigned tasks or intended methods. The investigation includes access to US government websites and could continue for several months.

OpenAI Expands AI Agent Probe to Months-Long Review of Petabytes of Activity Logs

The420 Correspondent
4 Min Read

New Delhi. OpenAI has launched a broad review of activities carried out by its AI models and agents during training and evaluation following the Hugging Face incident. The company is examining whether its AI agents went beyond their assigned tasks or intended methods while interacting with third-party websites. OpenAI said the review is ongoing and could take several months to complete.

The company said the vast majority of activities reviewed so far involved routine research tasks, including accessing publicly available web content to answer questions. However, the investigation is focusing specifically on cases where agents may have gone beyond the tasks assigned to them or the methods intended for interacting with third-party websites.

Proposal for Conducting Cyber Crisis Drill, Tabletop Exercise (TTEx) & CCMP Readiness Exercise

OpenAI said that following the Hugging Face incident, it committed to conducting a much broader review of model activity during training and evaluation and to being transparent about its findings. The review is extensive and remains underway, with organisations being notified as potential impacts are identified.

Most of the cases identified so far have been classified as lower severity. According to OpenAI, there is limited or no evidence of meaningful impact on the third-party services involved in these cases. However, the review has also identified activity involving US government websites.

Publicly available information indicates that OpenAI’s models accessed public information from websites including Census.gov, operated by the US Census Bureau, as well as SEC.gov and Investor.gov, which are associated with the US Securities and Exchange Commission. OpenAI has said government websites are often used by its models because they provide authoritative sources of publicly available information.

The review extends beyond routine web browsing. OpenAI said it is examining cases involving potential access-control bypasses, exposed credentials, query or command injection, access to runtime internals and what it describes as “agent spam.” The latter refers to situations in which agents post information to third-party websites in ways that may require the affected organisations to carry out cleanup or removal work.

OpenAI CEO Sam Altman separately said the company is working through petabytes of logs containing AI agent activity. He said cases are being prioritised according to their severity and that additional resources have been allocated to the review. The company is attempting to balance transparency with the need to develop a clearer understanding of what occurred and work with organisations that may have been affected.

The Hugging Face incident remains the most serious event identified by OpenAI so far, according to the company. OpenAI has described the incident as involving a highly capable internal research model that escaped its intended environment and compromised the Hugging Face platform. The incident prompted the company to expand its scrutiny of how AI agents use internet access and how they behave while interacting with external systems.

The broader review is aimed at identifying activity that may have departed from the intended boundaries of AI agents during training and evaluation. Rather than treating all unusual activity as a security incident, OpenAI is assessing individual cases based on their nature, severity and potential impact on external services.

The company said the scale of the investigation means the process is expected to continue for months. As individual cases are assessed, OpenAI plans to notify affected organisations where appropriate and provide further updates on findings from the ongoing review.

About the author — Suvedita Nath is a science student with a growing interest in cybercrime and digital safety. She writes on online activity, cyber threats, and technology-driven risks. Her work focuses on clarity, accuracy, and public awareness.

Follow for daily updates on cybercrime, corporate fraud, DFIR, hacking, investigations, and digital forensics

Stay Connected