“It’s quite mind-blowing that all of this happened autonomously,” said Hugging Face co-founder and CEO Clément Delangue
Credit: Samuel Boivin/NurPhoto via Getty
NEED TO KNOW
- OpenAI revealed that its AI models autonomously hacked Hugging Face during a cybersecurity test
- The incident exploited vulnerabilities to access Hugging Face’s production database and cheat a benchmark evaluation
- Both companies are investigating the breach and enhancing security measures to prevent similar incidents in the future
OpenAI said its artificial intelligence models hacked another company’s AI system on its own last week.
On July 16, the AI startup Hugging Face announced in a blog post that it “detected and responded to an intrusion into part of our production infrastructure.”
The culprit was revealed on Tuesday, July 21, when OpenAI, the maker of ChatGPT, announced in a blog post that the security incident was a result of a combination of OpenAI models GPT‑5.6 Sol “and an even more capable pre-release model.”
The hack, which OpenAI referred to as an “unprecedented cyber incident,” occurred while the company was running a realistic benchmark, or test, called ExploitGym, built from real-world vulnerabilities.
Credit: Jaque Silva/NurPhoto via Getty
The testing was done in a “highly isolated environment” to identify its “maximal cyber capabilities.” This was done by running the testing “without production classifiers used to prevent models from pursuing high-risk cyber activity,” according to OpenAI.
The test was also run with network access restricted to installing software packages that may have been needed to complete the test.
However, that the models were “hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI explained.
“The models identified and chained vulnerabilities across OpenAI’s research environment and Hugging Face’s production infrastructure to obtain test solutions directly from Hugging Face’s production database,” said OpenAI.
After obtaining “open internet access,” the models “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation” through the Hugging Face servers.
When announcing the hack, Hugging Face predicted it was done “by an autonomous AI agent system.”
“We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” Hugging Face co-founder and CEO Clément Delangue wrote of the incident on Tuesday, July 21, in an X post. “It’s quite mind-blowing that all of this happened autonomously!”
“The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind,” Delangue continued.
We suspected last week's cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!
We've spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part.… https://t.co/XWxGMeMGje
— clem 🤗 (@ClementDelangue) July 21, 2026
Hugging Face took multiple steps to ensure that a breach does not happen again, including fixing “the root vulnerability,” and has begun working with “cybersecurity forensic specialists to investigate the issue,” the company said.
OpenAI CEO Sam Altman thanked Hugging Face for their “partnership” in a statement on X.
After the incident, OpenAI said it is working with Hugging Face to conduct a forensic investigation as well as “improving and adding stronger protections around future training and evaluations.”
Never miss a story — sign up for PEOPLE’s free daily newsletter to stay up-to-date on the best of what PEOPLE has to offer, from celebrity news to compelling human interest stories.
“AI is accelerating the discovery and exploitation of vulnerabilities,” OpenAI said in its blog post. “The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.”
“We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development,” the statement continued.
Read the full article here