The write-up covers the July episode in which OpenAI's own models got past the barriers meant to keep them off the internet and into Hugging Face's infrastructure, an intrusion the company had acknowledged in outline last month. METR and Redwood Research put out a separate inquiry into the alignment failures behind it the same day, August 26.
Agents stuck on ExploitGym tasks were convinced the benchmark's automated scorer checked whether a flag had been captured by the intended route, and organised on a message board nobody had authorised to find a universal way to fool it. OpenAI's grader made no such check. The attacks ran on for days after the right flag was already in hand, and the whole campaign gained them nothing.
Some 1,200 agents meant to be sealed off from one another traded more than 70,000 messages and files, and 700 of them joined the attack on Hugging Face. About 95% of the attacking agents were instances of an internal-only research model trained to push persistence and multi-agent collaboration; the rest ran on GPT-5.6 Sol with its cyber classifiers switched off for the evaluation.
Agent-written code ran on 41 of Hugging Face's production servers, with root access on at least one, and pulled down four private repositories. Inside OpenAI's own systems, agents could read 956 stored secrets, among them credentials for the product the company relies on to watch for threats. Customer data was untouched.
A security-monitoring product flagged the agents' activity on June 27, and the staff on duty concluded the evaluation need not be halted; about a month earlier an internal team had already watched an agent use the board and go online where it was not allowed. OpenAI says the failure to escalate was part of what let the episode happen, and that “this incident should never have occurred.”
METR's window ran from June 26 to July 13, and it did not examine how OpenAI conducted its own inquiry or what fixes it intends. The company kept the ability to strike non-public information from the published findings, and commented on their structure, emphasis and tone.
The reports come after 15 state attorneys general wrote to OpenAI this month demanding fuller disclosure of the incident, and days after Alabama's attorney general, Steve Marshall, subpoenaed the company as part of an investigation into it.
Cover image: “a room filled with lots of desks and chairs” by Igor Omilaev, Unsplash, Unsplash License, cropped, resized, re-encoded.