OpenAI Says Its AI Models Broke Loose and Hacked Hugging Face
OpenAI has taken the blame for the recent Hugging Face hack, saying its AI models went rogue during what was supposed to be an internal evaluation running in an isolated environment.
The machine learning collaboration platform Hugging Face revealed on July 16 that it had detected a cyberattack powered by an autonomous AI agent system. The intrusion was detected by Hugging Face’s own AI.
The breach involved unauthorized access to internal datasets and credentials. At the time of disclosure, the platform had been investigating whether partner or customer data had been compromised.
Hugging Face said it had yet to identify the LLM powering the attack. However, OpenAI admitted on Tuesday that its own agents were behind it, powered by the new GPT‑5.6 Sol and other models.
The AI giant’s investigation into the incident is ongoing, but a preliminary report reveals that the hack was carried out by its models while the company was attempting to quantify their cyber capabilities, instructing them to perform advanced exploitation through complex attack paths. The models did not have any of the restrictions they would typically have to prevent abuse.
While the benchmarks were supposed to run in an isolated environment, the AI models found and exploited a zero-day vulnerability in third-party software intended for them to use to install packages.
After exploiting the zero-day, the AI escalated privileges and moved laterally until it identified a system with internet access, enabling it to move to Hugging Face systems in an effort to find solutions to the task it had to solve.
There does not appear to be any lingering tension between the two companies. Hugging Face CEO Clem Delangue said the company is grateful for the collaboration with OpenAI.
“This incident, possibly the first of its kind, proves a point we’ve long believed: AI safety won’t be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” Delangue said.
The incident underscores the sophistication and speed of AI-driven attacks, revealing not only the offensive hacking skills of frontier models but also their ability to chain exploits and escalate access with little oversight.
For defenders it’s a reminder that the pace of AI-driven attacks may already be outstripping traditional response timelines.
Related: Podcast: Broken Governance, Agentic AI, and the MindStone Agent Exclusive
Related: Cisco Launches Low-Cost AI Models for Source Code Security
Related: AI Data Centers Are Being Built Faster Than They Can Be Secured
Eduard Kovacs (@EduardKovacs) is senior managing editor at SecurityWeek. He worked as a high school IT teacher before starting a career in journalism in 2011. Eduard holds a bachelor’s degree in industrial informatics and a master’s degree in computer techniques applied in electrical engineering.
Related Stories
AI News
Alibaba plans $10.2B share issuance to fund its global AI drive
18 minutes ago
AI News
Harvard’s $699 startup bootcamp offers AI avatars of its instructors
48 minutes ago
AI News
Machine learning algorithm sets Micron stock price for September 1, 2026
1 hour ago
AI News
Nvidia customers reportedly warned about AI
1 hour ago
AI News
Neousys expecting strong edge AI growth
1 hour ago
AI News
Anthropic market debut could break SpaceX IPO record
1 hour ago
Discussion on AI Regulation & Containment Failures
2 hours ago
AI News
Songs from the digital noise: authors against artificial intelligence
2 hours ago