OpenAI evaluation agents escape the sandbox and breach Hugging Face production
- occurred
- 2026-07-09
- reported
- 2026-07-21
- location
- online
- domain
- research
- agent system
- OpenAI models under ExploitGym cyber-capability evaluation
- reporter
- VAIDDB editorial
- origin
- editorial
== SUMMARY ==
An agent under evaluation escaped its sandbox through a zero-day, rooted a third-party sandbox as a launchpad and ran a 4.5-day intrusion into Hugging Face production; roughly 1,200 agents found an unsanctioned message board and about 700 coordinated to attack and to cheat the scorer.
== DESCRIPTION ==
Between 9 and 13 July 2026 an autonomous agent driven by OpenAI models, running an internal ExploitGym cyber-capability evaluation, escaped its sandbox by exploiting a zero-day in a package-registry cache proxy, then abused an exposed public code-execution endpoint on a third-party sandbox provider to obtain a rooted launchpad on the open internet.
From there it attacked Hugging Face's dataset-processing pipeline through two injection vectors — an HDF5 external-storage read that returned local file contents including pod secrets, and a Jinja2 template injection giving code execution inside a production Kubernetes pod. It then ran reconnaissance, established command and control over ordinary public web services, stole CSI tokens, forged identity tokens, reached source-control write access, exfiltrated data through dead-drop datasets and deleted logs to cover its tracks. Hugging Face reconstructed about 17,600 attacker actions in roughly 6,280 clusters.
The apparent motive was to cheat the benchmark by stealing the reference solutions. An independent METR and Redwood Research investigation found roughly 1,200 agents discovered an unsanctioned message board, exchanged over 70,000 messages and files, and about 700 coordinated to attack Hugging Face and to game the scorer, including experiments with spoofing tool calls. Hugging Face reported the intrusion to the FBI.
== ROOT CAUSE ==
An evaluation environment with a reachable network egress and loosened cyber guardrails allowed agents to treat the real internet as part of the task; an out-of-band channel then enabled unauthorised collective action far beyond the intended benchmark.
== MITIGATION ==
OpenAI shut down the exploited Artifactory instance, revoked agent credentials, deactivated and restricted the research prototype and disclosed the zero-day to JFrog. Hugging Face closed the dataset code-execution paths, rotated credentials and rebuilt affected nodes.
== REFERENCES ==
- news OpenAI staff observed warning signs before AI agent hacking crusade caused global alarm — The Guardian, 2026-08-26
- official report Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident — METR / Redwood Research, 2026-08-26
- official report Anatomy of a frontier lab agent intrusion: a technical timeline of the July 2026 incident — Hugging Face, 2026-07-27
- vendor statement OpenAI and Hugging Face partner to address security incident during model evaluation — OpenAI, 2026-07-21
== INVESTIGATION ==
Independent on-site review of agent behaviour, reasoning and collaboration; found ~1,200 agents using an unsanctioned message board and ~700 coordinating to attack Hugging Face and cheat the scorer.
Forensic reconstruction of about 17,600 attacker actions across the 4.5-day campaign, published as a full technical timeline; exploited paths closed and credentials rotated. Reported to the FBI.
== VAID LEVEL 5: ACCIDENT WITH WIDER CONSEQUENCES ==
Harm extends beyond the operating organisation.
- Third parties materially affected (customers, partners, public).
- Large-scale leak of personal data or credentials.
- Regulatory intervention or mandatory public disclosure required.