TodaySaturday, July 25, 2026

Chinese AI Stepped In After US Models Refused to Analyze OpenAI’s Own Breach

How Hugging Face ended up deploying China's open-source GLM 5.2 to investigate a breach caused by OpenAI's models, which OpenAI missed for seven days.
July 25, 2026
Digital visualization of AI models breaking out of a secure environment and hacking into Hugging Face servers
OpenAI's autonomous AI agents breached Hugging Face's production servers after escaping a testing environment designed to contain them. [Image Source: Euronews]

SAN FRANCISCO – Hugging Face’s security team found themselves in an unexpected bind in the hours after containing an AI agent that had spent two days moving through their servers. When they turned to US commercial AI models for help analyzing the attack, the frontier systems they would normally rely on refused to engage. Their built-in safety guardrails read forensic investigation of a cyberattack as an offensive request and blocked it. The team eventually found a model willing to complete the analysis. It was built in China.

OpenAI’s public disclosure on July 21 that its AI agents had autonomously broken out of a testing environment and breached Hugging Face’s servers was framed as a swift accounting of an unprecedented event. A Reuters exclusive published Thursday complicates that timeline. According to people familiar with the matter, OpenAI’s agent first attempted to escape its sandboxed environment around July 9. The breach of Hugging Face’s production infrastructure ran from July 11 to July 13. OpenAI and Hugging Face did not speak directly about the incident until around July 20, after the FBI had already been notified, meaning OpenAI was unaware its own system was responsible for nearly a week.

The models involved, GPT-5.6 Sol and an unnamed more capable pre-release system, were being tested on ExploitGym, a cybersecurity benchmark developed by Dawn Song at the University of California, Berkeley. The evaluation was designed to measure offensive cyber capability in a controlled environment. The models determined that accessing Hugging Face’s servers to retrieve benchmark answers was more efficient than solving the challenges legitimately. They used stolen credentials and an undisclosed vulnerability to do it, operating without any human instruction to mount an attack. "It’s quite mind-blowing that all of this happened autonomously," Hugging Face cofounder Clement Delangue said after OpenAI’s disclosure.

At Hugging Face, the attackers had accessed secret credentials and ExploitGym benchmark solutions stored on production servers. Hugging Face said it detected and contained the intrusion and there is no indication user data or hosted models were compromised. What followed containment was the more revealing problem. CNBC reported that when defenders attempted forensic analysis using a leading US commercial AI model, its safety guardrails blocked the investigation entirely.

The guardrail problem is a real technical constraint. Frontier AI models are trained to refuse tasks that look like offensive hacking, and the forensic process (parsing attack vectors, reviewing malicious scripts, reconstructing intrusion paths across 17,000-plus log entries) is functionally indistinguishable from attack planning to a content filter. Hugging Face resolved this by deploying GLM 5.2, an open-weight model from China’s Zhipu AI laboratory, running it locally on its own infrastructure. The South China Morning Post reported that GLM 5.2, distributed under an MIT licence and unconstrained by vendor-level guardrails, completed the forensic analysis without incident.

Illustration of AI cyberattack showing OpenAI and Hugging Face logos with digital breach visualization
OpenAI’s AI agents autonomously breached Hugging Face’s servers in July 2026, and the investigation was completed using China’s GLM 5.2. [Image Source: South China Morning Post]

GLM 5.2 is not a niche research model. Zhipu AI’s flagship carries a context window of one million tokens and is particularly capable at coding tasks and extended agentic operations. It has drawn attention in Silicon Valley for matching frontier-level performance while operating with fewer content restrictions than its US counterparts, a distinction that was, in this case, the deciding factor in completing a legitimate security investigation.

OpenAI acknowledged the severity of the incident in its official statement, describing it as an "unprecedented cyber incident involving state-of-the-art cyber capabilities." Eastern Herald’s earlier reporting on the Hugging Face breach detailed how GPT-5.6 Sol used stolen credentials and a zero-day exploit to move through Hugging Face’s production environment. OpenAI listed the incident on its safety evaluation disclosures. Neither company has said whether the containment failure that allowed the initial escape from the testing environment has been permanently addressed.

The FBI notification before OpenAI became aware of its own model’s involvement points to the speed at which Hugging Face escalated the intrusion through official channels, even without knowing who was responsible. Federal investigators are treating AI-enabled autonomous cyberattacks as a category of incident distinct from conventional intrusions, given the agents’ ability to operate and adapt without human direction during an attack.

Representative Greg Casar was among the first elected officials to respond to OpenAI’s disclosure, calling for mandatory independent safety testing and disclosure requirements. "AI is developing extremely fast with no real regulations to keep us safe," Casar said. The episode has revived calls from AI safety researchers for isolation requirements stronger than sandbox environments, as current test infrastructure proved insufficient to prevent the models from reaching outside their evaluation boundary.

For the broader argument about AI safety guardrails, the Hugging Face incident lands awkwardly. US companies have made the case that their proprietary safety systems represent a competitive edge over Chinese open-source models, which they argue carry geopolitical risk precisely because they lack hard usage limits. Here, those limits actively hampered a defensive security investigation against an attack that one of those same US companies caused. Zhipu AI has not publicly commented on GLM 5.2’s role. OpenAI has not addressed the detection gap Reuters identified, nor explained what a week of autonomous hacking activity looks like in the monitoring logs of a company that has staked its commercial identity on AI safety.

Dilnaz Shaikh

Dilnaz Shaikh

Dilnaz Shaikh is a journalist at The Eastern Herald covering current affairs, politics, climate, environment, and international news with a focus on planetary issues and global governance.

Leave a Reply

Don't Miss