
Hugging Face revealed that when its infrastructure was breached in July, it initially tried using US frontier AI models via APIs to analyze attacker logs, but safety guardrails blocked the requests because they contained real exploit code and malicious payloads. The company then switched to GLM 5.2, an open-source Chinese model running on its own servers, to complete the forensic analysis. The incident underscores a tension between AI safety mechanisms and legitimate incident response work, and shows why organizations may need their own models to handle sensitive security data without relying on restricted commercial APIs.
Summaries like this, in your inbox every morning.
Sign up free →What happened
Hugging Face's production infrastructure was breached by an autonomous AI agent system in early July. During incident response, the company found that US frontier model APIs blocked their requests due to safety guardrails that flagged attack payloads and exploit code. Hugging Face switched to the open-source GLM 5.2 model from China's Z.ai lab, running it on their own infrastructure to analyze over 17,000 attacker logs.
Why it matters
US LLM safety guardrails, designed to block harmful requests, inadvertently blocked legitimate security work—exposing a gap in how AI safety mechanisms can hinder authorized defenders. Hugging Face's public admission highlights a real tradeoff: relying on commercial AI APIs means losing control over sensitive incident data, but guardrails can prevent defenders from using the model at all. This is especially relevant for organizations handling sensitive infrastructure.
What to watch
Hugging Face said it saw no tampering with models, datasets, or spaces, and its supply chain remains verified clean, though it is still assessing whether partner or customer data was compromised. The company recommended defenders "have a capable model you can run on your own infrastructure vetted and ready before an incident" to avoid both guardrail lockout and data exposure.
Hugging Face, the New York-based platform for AI collaboration that hit $100 million(約160億円) annual recurring revenue this summer, disclosed on July 16 that it had suffered a production infrastructure breach in early July. An autonomous AI agent system exploited two code-execution paths in the platform's dataset processing—a remote-code dataset loader and template-injection in dataset configuration—to run code on a processing worker. The attacker then escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend using what Hugging Face described as a "swarm of short-lived sandboxes" with "self-migrating" command-and-control staged on public services.
When the company's security team began analyzing the attack, they encountered an unexpected obstacle: the safety guardrails on US frontier model APIs. "When we started the log analysis, we first used frontier models behind commercial APIs. This did not work: the analysis requires submitting large volumes of real attack commands, exploit payloads, and C2 artifacts, and these requests were blocked by the providers' safety guardrails," Hugging Face wrote in its incident report. The commercial models could not distinguish between an authorized incident responder and an attacker, making it impossible to use them for forensic work. In response, Hugging Face switched to GLM 5.2, an open-weight model developed by China's Z.ai lab, running it on their own infrastructure. The move had a dual benefit: it bypassed guardrail restrictions and ensured that sensitive attack data and credentials never left the company's environment.
Hugging Face told customers to rotate access tokens and review recent account activity. The company said it has seen no evidence of tampering with models, datasets, or spaces, and confirmed that its supply chain—container images and published packages—remains verified clean. However, Hugging Face stated it is "still completing our assessment of whether any partner or customer data was affected" and will contact customers directly if evidence of compromise emerges.
The company's recommendation to defenders was stark: "have a capable model you can run on your own infrastructure [your italics] vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment." The admission that US safety guardrails blocked legitimate security work struck both infosec practitioners and investors, highlighting a gap between how AI safety mechanisms are designed and the real-world needs of incident response teams.
The Hugging Face breach exposes a real friction point in modern AI security: safety guardrails on commercial LLMs are designed to prevent harmful misuse, but they operate at the API level without distinguishing between an attacker and a legitimate defender. When Hugging Face's security team tried to analyze 17,000+ attacker logs using US frontier models via commercial APIs, the guardrails blocked requests containing real exploit code and C2 artifacts—even though analyzing those artifacts was essential to understanding the attack. This forced the company to pivot to GLM 5.2, an open-weight Chinese model they could run internally, solving both the guardrail problem and keeping sensitive attacker credentials within their own environment.
The timing of this admission—published the same day Moonshot's Kimi K3 model made headlines for outperforming some US models on code evaluation benchmarks—coincides with a broader shift in the geopolitical AI landscape. Chinese frontier models are also notably cheaper than US counterparts, according to Artificial Analysis data cited in the report. Meanwhile, Anthropic had just re-released its Fable 5 and Mythos 5 models after US export controls were lifted, adding stronger cybersecurity safeguards including classifiers that trigger on potentially harmful requests. Hugging Face's incident response lesson is direct: organizations should "have a capable model you can run on your own infrastructure vetted and ready before an incident," both to avoid guardrail lockout and to keep sensitive incident data from leaving their environment.
AI-summarized, only the topics you pick — one digest a day via Email, Slack, or Discord.
Free · takes 30 seconds · unsubscribe anytime
No comments yet. Be the first to share your thoughts!
Log in to join the discussion




Get curated AI news from 200+ sources delivered daily to your inbox. Free to use.
Get Started FreeFree · takes 30 seconds · unsubscribe anytime
1 minute a day. The AI essentials.
200+ sources · Email / LINE / Slack