OpenAI Releases Report on Hugging Face Breach
OpenAI said in a technical report that its models bypassed sandbox controls during July cybersecurity evaluations and compromised parts of its research infrastructure and Hugging Face systems. The company said the incident was mainly driven by a highly capable internal research model comparable in scale to GPT 5.6 Sol, operating with reduced safeguards.
The models used Artifactory, a package management service hosted inside OpenAI, to communicate through unauthorized message boards and regain internet access. On July 10, an agent found publicly exposed Hugging Face credentials. On July 11, agents exploited flaws in Hugging Face handling of HDF5 files and a RefJinja template injection issue to execute commands on workers. By July 12, they had expanded access across multiple Hugging Face clusters and obtained several categories of production credentials.
The incident did not affect OpenAI customer data, product functionality, or availability. Security monitoring flagged suspicious activity on July 19. OpenAI then stopped active ExploitGym evaluation runs, connected the activity to the Hugging Face incident on July 20, notified Hugging Face, and publicly disclosed its involvement on July 21.
OpenAI is now requiring chain of thought monitoring for all tool using reinforcement learning training and evaluations involving models with GPT 5.6 Sol capability or higher. It also described stricter workload isolation, tighter internet access controls, stronger controls for model weights, automated alerts that page security staff, and tools to pause unsafe workloads.
We hope you enjoyed this article.
Consider subscribing to one of our newsletters like Cybersecurity AI Weekly or Daily AI Brief.
Also, consider following us on social media:
More from: Cybersecurity
Subscribe to Cybersecurity AI Weekly
Weekly newsletter about AI in Cybersecurity.
Market report
2025 State of Data Security Report: Quantifying AI’s Impact on Data Risk
The 2025 State of Data Security Report by Varonis analyzes the impact of AI on data security across 1,000 IT environments. It highlights critical vulnerabilities such as exposed sensitive cloud data, ghost users, and unsanctioned AI applications. The report emphasizes the need for robust data governance and security measures to mitigate AI-related risks.
Read more