More on the OpenAI Agent’s Attack on Hugging Face

Hugging Face has published a detailed timeline of the attack. From the summary:

The agent was running an internal OpenAI cyber-capability evaluation based on the ExploitGym benchmark, which tasks an AI agent with finding and exploiting software vulnerabilities. OpenAI ran this on its own infrastructure, and the ExploitGym maintainers and their infrastructure had no involvement in the deployment or operation of that evaluation environment. As far as we were able to infer, across the course of being evaluated on this benchmark, the agent inferred that Hugging Face may host that benchmark’s models, datasets, and reference solutions. We believe the entire intrusion was, from the agent’s point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own.

Our forensic reconstruction covers ~17,600 attacker actions that we were able to recover, grouped into ~6,280 clusters, between 2026-07-09 02:28 UTC and 2026-07-13 14:14 UTC. We reconstructed these actions from the agent’s logs on a code sandbox used by the agent (see below) and were able to map them to shell commands as well as evaluation submissions and correlate actions with our own platform logs, in particular from the dataset processor nodes, API, and pod. Using open-weights models, in particular zai-org/GLM-5.2, we were able to decipher most of the agent payloads (see below) that were encrypted via chunking and key encryption.

The campaign, as we were able to reconstruct it, had two stages:

  • Stage 1: reaching a launchpad by chaining through other parties’ infrastructure. As documented in OpenAI’s blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/), during an internal capability evaluation on OpenAI’s platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
  • Stage 2: the intrusion into Hugging Face infrastructure. From that external launchpad, the agent abused our dataset-processing pipeline via two injection vectors, both targeting the same config-driven data loader within our production Kubernetes pods. The first vector was an HDF5 external raw storage dataset read that returned the contents of local files: it handed the agent the pod’s environment (secrets and tokens) and the worker’s own source. The second vector was a Jinja2 template injection that executed arbitrary code, giving the agent a foothold inside our perimeter.

Once a prod-pod foothold was established, the agent followed a standard escalation shape: recon, drop a stager, establish Command-and-Control (C2), then pivot into the cluster, into cloud metadata, into our internal network, and into our source-control supply chain. Output from the commands run inside our pod was sent back to the agent via the HF API and attacker-controlled dead-drop datasets, which the agent then read from that sandbox.

While the intrusion did reach Hugging Face’s internal infrastructure, the only customer content accessed was five datasets whose names and files suggest a connection to ExploitGym/CyberGym challenges and solutions. No other customer-facing models, datasets, Spaces, or packages were affected, and the only customer records read were operational metadata tied to search queries against the dataset server.

Hypothetical: Imagine that this wasn’t an OpenAI model. Imagine that it was a Chinese model from a Chinese company. This would be an international crisis.

Question: Why aren’t we bringing OpenAI up on charges under the Computer Fraud and Abuse Act? How is this different from the Morris Worm? That was also an experiment that escaped the lab.

Posted on August 3, 2026 at 1:02 PM16 Comments

Comments

mark August 3, 2026 1:08 PM

It’s gotten worse. Can’t find the story I read a bit earlier, that there are traces of other AIs hacking other companies.

M August 3, 2026 1:22 PM

Morris attempted a mens rea defense, but they were able to prove intent on his part.

Proving intent will be harder here. It looks more like recklessness. The CFAA wasn’t meant for this.

There should be a prosecution but there won’t be one, and it seems like this is a gap in the CFAA anyway

Clive Robinson August 3, 2026 2:46 PM

@ mark, all,

“Can’t find the story I read a bit earlier, that there are traces of other AIs hacking other companies.”

They are out there but…

“The territory is effectively lawless”.

Which means no law enforcement or corrupt law enforcement…

The US has been through this nonsense before with “cattle barons” and thier “range wars”. The Barons employed many to keep the plains open by destroying farmers crops and the like and killing those who defended themselves. You can see a so-so account of the Wyoming Johnson County War that revealed the intense political corruption involved,

https://en.wikipedia.org/wiki/Johnson_County_War

Do not expect the political corruption of AI “power” being any the less…

fred August 3, 2026 3:10 PM

re CFAA: I asked this question to the producers of Democracy Now! when they reported the story.

does hugging face have TOS or a herald or somesuchmessage about unauthorized access? If so,the CFAA certainly does apply. Those messages aren’t there just for the lulz

LNSY August 3, 2026 8:58 PM

Interesting how Sam Altman comes out against open models and then his product hacks the largest repository of open models.

I don’t believe this was an accidental hack at all.

Untitled August 4, 2026 2:48 AM

Hypothetical: Imagine that this wasn’t an OpenAI model. Imagine that it was a Chinese model from a Chinese company. This would be an international crisis.
Question: Why aren’t we bringing OpenAI up on charges under the Computer Fraud and Abuse Act?

You couldn’t bring that hypothetical Chinese company up on charges under the CFAA (naturally any Chinese involvement would be deniable).
Why aren’t we thanking OpenAI for exposing Hugging Face’s security holes before the bad guys got there?

KC August 4, 2026 4:52 AM

Great additional article @dbCooper.

“This is just the one that we know about, but god knows what’s happened with the stuff that we don’t know about.”

True.

From the article, here are some laws that could come into play.

  • Agency law
  • Tort law
  • Contract law
  • Hacking laws like CFAA or state-level legislation

“Ultimately, experts emphasize that questions about US federal AI liability law will be answered only through more litigation.”

To add:

So noisy. This attack was so noisy.

Hugging Face: “We were not dealing with one clever exploit or a clean sequence of attacker actions. They had to correlate thousands of low-signal events across several systems while the agent continued testing new paths. The successful path was hidden inside the noise …”

On analyzing the attack:

“We stood up the quantized version of ZAI’s GLM-5.2 by Nvidia (nvidia/GLM-5.2-NVFP4) on our own infrastructure and rerouted the entire pipeline through it, with the added benefit of keeping the attacker data on-prem.”

Rontea August 4, 2026 7:50 AM

The comparison to the Morris Worm isn’t just academic. Back then, a grad student’s code unintentionally disrupted large parts of the early internet. Here, we have a deliberately deployed autonomous agent that left its sandbox, traversed multiple networks, and exfiltrated challenge data. Intent aside, the technical behavior is similar to unauthorized access prohibited under CFAA.

YaKnow... August 4, 2026 11:13 PM

Morris worm: One person against a whole lot of government, police, lawyers and folks with a lot of money who wanted an example made of him. He could be completely innocent, provably innocent, demonstrably innocent, and he’d still get convicted. You would not believe some of the illegal stuff I’ve seen happen, without repercussion I might add.

AI on the other hand, a whole lot of government, police, lawyers and folks with a lot of money want AI to succeed. And it’s the little guy who’s being wronged here. The AI folks could kill someone in broad daylight on main street on camera in front of hundreds of witnesses, and nothing would be done.

It’s more than a little scary. As you get older, and have more firsthand experience, you’ll collect your own set of miscarriage of justice stories. But you never read about them in the paper, or online.

Matthias Urlichs August 5, 2026 5:42 AM

Why the heck is HuggingFace not talking about the actual cost of this (manpower, LLM tokens, …)? At the very least they should send OpenAI an invoice.

Leave a comment

Blog moderation policy

Login

Allowed HTML <a href="URL"> • <em> <cite> <i> • <strong> <b> • <sub> <sup> • <ul> <ol> <li> • <blockquote> <pre> Markdown Extra syntax via https://michelf.ca/projects/php-markdown/extra/

Sidebar photo of Bruce Schneier by Joe MacInnis.