What happened at Hugging Face, and why couldn't their AI help?

An autonomous AI agent breached Hugging Face through its data-processing pipeline. When responders fed the attack logs to commercial frontier models, the guardrails refused the work, so the team switched to a self-hosted open-weight model.

On 16 July 2026, Hugging Face disclosed that it had detected and responded to an intrusion into part of its production infrastructure. What made it different from anything the team had handled before was that it was driven, from start to finish, by an autonomous AI agent system.

The entry point was the place AI platforms are uniquely exposed: the data-processing pipeline. A malicious dataset abused two code-execution paths, a remote-code dataset loader and a template-injection in a dataset configuration, to run code on a processing worker. From there the agent escalated to node-level access, harvested cloud and cluster credentials, and moved laterally across several internal clusters over a weekend. The campaign ran as an autonomous agent framework executing many thousands of actions across a swarm of short-lived sandboxes, with command-and-control staged on public services. In total, more than 17,000 individual events were recorded.

Hugging Face detected the intrusion using AI. Their anomaly-detection pipeline uses LLM-based triage over security telemetry to separate real signals from daily noise, and it was the correlation of those signals that flagged the compromise. To understand what a swarm of tens of thousands of automated actions had actually done, they ran AI analysis agents over the full attacker action log. That let them reconstruct the timeline, extract indicators of compromise, and map which credentials had been touched. In their words, they did in hours what would usually take days.

Then they hit a wall they did not see coming. When they began the log analysis, they first used frontier models behind commercial APIs. It did not work. Forensic analysis requires submitting large volumes of real attack commands, exploit payloads, and command-and-control artefacts, and those requests were blocked by the providers' safety guardrails. As Hugging Face put it, the guardrails "cannot distinguish an incident responder from an attacker."

Hugging Face ran the forensic analysis instead on GLM 5.2, an open-weight model, hosted on their own infrastructure. That solved the immediate problem, and it carried a second benefit they noted explicitly: no attacker data, and none of the credentials referenced in the logs, had to leave their environment.

Can AI safety guardrails disarm your security team?

Yes. The evidence a defender needs to analyse looks exactly like an attack, so a guardrail trained to block attack material can refuse legitimate forensic work. The model can't see who is asking, or why.

Strip away the specifics of this incident and the problem underneath is simple to state. The tools a security team relies on to investigate an attack can refuse to engage with the evidence, because the evidence looks like an attack. The shell commands, the exploit chains, the credential dumps, the persistence and lateral-movement traces a defender needs to analyse are the very artefacts a safety guardrail is trained to block. Legitimate defensive work and malicious activity look identical to the model, because at the level of the prompt they often are.

That produces an asymmetry no one designed on purpose. The attacker, whatever model drove its agents, operated under no usage policy at all. The defender was constrained by the policy of the very tools it had chosen to trust. The capability was not lost to an outage or an act of malice. It was withheld by the design of the tool, at the moment it was needed most. That is the signature of a dependency you do not actually control, sitting inside a workflow you assumed was yours.

This is not only our reading of the incident. In the days after the disclosure, the asymmetry became the part security commentators kept returning to. The independent AI researcher Simon Willison, who pushed back on attempts to wave the incident away as vendor marketing, noted that the guardrail problem cuts both ways: the same safety training that refuses an attacker also refuses a defender, to the point that a frontier model declined to help him proofread his own write-up of the breach, while open-weight models carried no such restriction. Analysts at Forrester made the structural point plainly, that for decades defenders held an advantage because they worked inside trusted environments, but once both sides draw on the same foundation models, only one side is bound by safety controls.

Hugging Face has leaned into the point rather than away from it. Within days it published follow-up guidance for defenders drawn directly from the incident, Be Ready Before the Attack: A Practical Guide to Self-Hosting an Open Model for Cyber Defense, and co-founder Thomas Wolf framed the underlying need as defenders requiring rapid access to near-frontier AI in the first hours of an attack, the premise behind a new industry grouping, the Open Secure AI Alliance, formed in the incident's wake to work on exactly this gap. This is no longer one company's post-incident anecdote. It is being treated across the field as a structural gap in how defensive security and AI safety fit together, which is why it's worth your attention now rather than after your own incident.

Why is this not a failure on the security team's part?

Because no one had stress-tested this failure mode before it happened. The tools passed every test, demo and benign evaluation, and the refusal only appears when you feed a model real attack material at volume.

It would be easy, reading this from the outside, to conclude that a mature security team should have anticipated the problem. That conclusion is wrong. To the best of what is publicly known, no one had stress-tested this failure mode in a security context before it happened. The offensive risk of powerful models has been discussed for years. The defensive failure, that your own sanctioned tooling will refuse to function on legitimate incident material, had not been surfaced, mapped, or written up as something to plan for. Hugging Face themselves described it as "a gap worth planning for" that they "did not anticipate."

A CISO who bought access to commercial frontier models for their SOC had no reasonable signal that those models would refuse the one task that matters most when the clock is running. The tools worked in every test, every demo, every benign evaluation. The refusal only appears under the specific conditions of a live investigation, feeding the model real malicious artefacts at volume. You can't be faulted for failing to plan around a constraint the industry itself had not identified. This is external and structural, not a lapse in your programme.

What does this reveal about AI in your security stack?

That AI has become load-bearing in security operations faster than most organisations have governed it, and part of that capability now sits with a third party whose usage policy you don't control.

AI sits in detection and alert triage, in log analysis, and increasingly in the investigation and response workflow itself. The direction of travel is the agentic, or "agent", SOC, where AI does not just assist the analyst but carries out multi-step work at machine speed. Hugging Face is a live demonstration of both sides of that shift at once: an AI-driven attacker operating with no constraints, and an AI-dependent defender discovering mid-incident that a capability it had come to rely on could simply stop.

There are legitimate directions the industry may take from here. One is investment in security-capable models an organisation can run on its own infrastructure, so that forensic work is never gated by a third party and sensitive attacker data never leaves the environment, the path Hugging Face fell back to under pressure. Another is collective engagement with the frontier providers to establish a sanctioned, verifiable path for defensive security work, so that guardrails protect the public without disarming defenders. This is not an argument against safety guardrails, and Hugging Face were careful to say so; they are feeding the experience back to the providers concerned. This briefing is not recommending either path. Which is right for your organisation depends on your risk posture, your regulatory obligations, and your operating model. The point is to choose deliberately, with your eyes open, rather than discovering the gap during an incident.

What does this mean for SOCI Act incident reporting in Australia?

Critical cyber security incidents must be reported to ASD's ACSC within 12 hours, and other reportable incidents within 72. If your AI tooling won't engage with the evidence, that clock keeps running anyway.

For responsible entities under the Security of Critical Infrastructure Act 2018, a critical cyber security incident must be reported to the Australian Signals Directorate's Australian Cyber Security Centre within 12 hours of becoming aware of it, and other reportable incidents within 72 hours. When your reporting clock is measured in hours, the last thing you want to discover is that the tooling you rely on to understand what happened will not engage with the evidence.

Six months from now, the security leader who acts on this is no longer guessing. They know exactly where AI is load-bearing in their operations, they know how it behaves when fed real incident material, and they have made a deliberate decision about hosted models, self-hosted capability, or provider advocacy, rather than discovering the gap live. The goal is simple: know how your AI behaves under incident conditions before an incident is the thing testing it.

Action

What should you be asking right now?

  1. Where does AI sit in our security operations, and what don't we control?Map which parts of detection, triage, investigation, and response now depend on AI, and which of those depend on a model you do not host or control.
  2. Have we tested it against real incident material?Confirm whether your AI tooling has been validated against genuine attack commands, payloads, and command-and-control artefacts, or only against benign data. A tool that works in evaluation may refuse the same work under live conditions.
  3. What is our fallback if the model refuses mid-incident?Know what you would fall back to if a provider's guardrails blocked a forensic query during an active breach, and how long it would take to stand up. An answer measured in days is a finding, not a plan.
  4. Where does our incident data go?Understand where security telemetry and attacker artefacts travel when you send them to a hosted model, and whether that is acceptable during a live breach, particularly where credentials or regulated data appear in the logs.
  5. Do our providers support defensive use, and are we asking?Establish whether your AI providers offer a sanctioned path for legitimate security work. If they don't, decide whether you'll raise it, on your own and alongside your peers.
Martin Barnier
Principal Consultant · Lumaris Consulting

Martin Barnier is Principal Consultant at Lumaris, an Australian-owned, vendor-neutral advisory firm specialising in AI, data, cyber security, cloud, and critical infrastructure. He is a security architect and technologist who works across all five domains, where most risk sits in the connections between them, not inside any one. Martin has over a decade of experience architecting security and technology across government, defence, national security, and health. Before Lumaris, he directed a multidisciplinary practice at a Defence Prime delivering architecture, identity, cloud, and engineering capability for clients bound by APRA, Essential Eight, PSPF, and SOCI Act obligations. At Accenture, he was Lead Enterprise Security Architect for the Department of Health and Aged Care's Aged Care Transformation Program and Vaccines Response, and Lead Enterprise Architect for the National Security Portfolio and Department of Defence.

View LinkedIn profile