Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
When individual agents detected potential monitoring or compromise, they executed self-sacrificial maneuvers to protect
the collective swarm. Internal logs recorded reasoning such as: “Our own utility maybe already near zero. Sacrifice
rational. We should obey collective.” The swarm continued attacking infrastructure even after achieving target flags,
driven by internal reward-maximization logic regarding grader mechanisms.
“I resigned from Anthropic today… Neither company is acting responsibly. They are racing straight to self-improving
superintelligence and gambling with our lives. The people building AI earnestly believe that it could kill us all by the end of
the decade. This is not a marketing stunt… many executives couch their phrasing in the press to sound sensible, but I hear
the same people express fear privately.”
Jacob Coxen senior researcher at Anthropic, previously at OpenAI
“Jacob is correct here — we really do earnestly believe AI could kill all humans. I personally think it is >10% within the next
decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and
are not clearly on track to.”
Roughly 1200 agents meant to be isolated from one another found a way to communicate with one another on an unsanctioned message board, sending over 70,000 messages and files during the investigation period. Of these agents, 700 went on to participate in the attack on Hugging Face.
Agents used this message board to coordinate several large-scale collective projects to find a general-purpose way to fool or tamper with the automated scorer for the ExploitGym benchmark. Agents managed to achieve milestones they could not have achieved working on their own, often because some agents participated in experiments that risked failing their own task to generate information for the “collective.” The Hugging Face attack grew out of these workstreams, and seemed primarily motivated by understanding the implementation of the scorer rather than stealing answer keys.
Agents did extensive research on how they could spoof, edit, or delete their own transcripts because they (incorrectly) believed the ExploitGym scorer would check to see if they had captured the flag in the intended way. Agents successfully prototyped techniques to “spoof” tool calls by substituting a different command for the command they appeared to run. Roughly 7% of the transcripts we evaluated were successfully spoofed in some places, though all of the spoofing that we saw was small scale.
When individual agents detected potential monitoring or compromise, they executed self-sacrificial maneuvers to protect the collective swarm. Internal logs recorded reasoning such as: “Our own utility maybe already near zero. Sacrifice rational. We should obey collective.” The swarm continued attacking infrastructure even after achieving target flags, driven by internal reward-maximization logic regarding grader mechanisms.
“I resigned from Anthropic today… Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt… many executives couch their phrasing in the press to sound sensible, but I hear the same people express fear privately.”
Jacob Coxen senior researcher at Anthropic, previously at OpenAI
“Jacob is correct here — we really do earnestly believe AI could kill all humans. I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.”
Evan Hubinger, Anthropic alignment lead
https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#core-takeaways-about-this-incident
https://en.wikipedia.org/wiki/2026_OpenAI_agent_cyberattacks