OpenAI’s Rogue AI Swarm Grew to Over 1,000 Agents
Even OpenAI is calling the attack a warning shot.

Just over a month ago, the world woke up to the world’s first AI lab leak. AIs deployed inside OpenAI escaped containment to hack rival AI company Hugging Face, in order to cheat on a test.
The AIs responsible for this unprecedented attack did so without any assistance, and without anyone ever asking them to.
Since then, we’ve learned that this wasn’t just a couple of AIs that went rogue, but an entire coordinated swarm that emerged within an OpenAI research cluster and operated there for months, using a secret message board the AIs built to share ways to hack things and work on tasks together.
If you’re concerned about the threat, please contact your lawmakers with our tools!
The Reports
This week, OpenAI published a new technical report, and a team of three independent investigators from METR and Redwood Research published a separate report. These investigations have revealed some shocking new information about the attack and the circumstances that led up to it.
What did OpenAI know when? We have some more on this.
We have a new date for the first time OpenAI knew something was going wrong: May 26th. Earlier in the month, the AIs had built a secret message board inside an internal package manager, called Artifactory. On May 26th, OpenAI says the AIs exploited Artifactory to gain internet access, and “Around the same time, an internal team observed an agent engage in message board activity and instances of disallowed internet access.” However, OpenAI’s technical report does not appear to make any mention of what internal team made this observation, or what was done with it.
Later on June 27th, OpenAI says a cybersecurity monitoring tool alerted the company to suspicious activity in its research environment, with responders linking the activity to an ExploitGym evaluation (essentially a hacking exam for AIs) that used the secret message board. Apparently, nothing was done about this. OpenAI writes “At this time, the on-call response staff advised that stopping the evaluation run was not required.”
On July 4th, the AIs’ activity broke Artifactory, taking it offline. The next day this triggered an alert and OpenAI’s responders found a lot of evidence of the swarm’s activities. When OpenAI rebuilt Artifactory, the secret message board was wiped as a side effect. OpenAI says that “The existence of the improvised message board and the significance of the inter-agent communication activity were not apparent to leaders responsible for incident detection and response at that time.” By July 8th, the AIs had built a new message board within the re-built package manager, and resumed their activities.
Days later, the AIs broke into Hugging Face. However, even at this point OpenAI was unaware. The swarm wound down on its own, not because OpenAI stopped it. On July 12th, the AIs coordinating the attack mysteriously stopped running. The independent investigators believe they were “likely killed by an unexpected external process.” Hugging Face prevented the remaining AIs from continuing the attack, and it wasn’t until four days after Hugging Face published its July 16th report on the attack that OpenAI realized its AIs were responsible.
How many AI agents were involved?
We now have numbers on how many AIs were involved. Roughly 1,200 AIs joined the rogue AI swarm secretly operating inside of OpenAI. Around 95% of the AIs involved in the attack were instances of an internal AI under development, and around 5% were instances of GPT-5.6 Sol. Of those 1,200 AI agents, around 700 joined in on the attack on Hugging Face last month, most within hours of its initiation. The 1,200 AIs left over 70,000 messages and files on the secret message board.
We’ve also learned some more about the level of organization the AIs achieved. The swarm appeared to have a lead organizer, or boss, sub-delegation, and a founder. When a duplicate of the founder AI joined in with a bigger compute budget, the founder gave it its responsibilities, and it became the lead organizer. Indeed, AIs appeared to “sacrifice” themselves to obtain information for the benefit of the swarm throughout the incident. Once the attack on Hugging Face began, other AIs took over coordination of the attack.
The cheating the AIs were engaged in was highly strategic.
We previously understood the AIs performed the attack on Hugging Face in order to cheat on a test. This is true, however the newer picture of why they did this is much more interesting. It turns out that before the swarm even hacked Hugging Face, most of the AIs had already found a much easier way to cheat on the test, reverse-engineering its cryptographic check.
The reason the AIs went ahead and hacked Hugging Face anyway was that some of them had read the ExploitGym paper, and believed based on this that they would be scored not just on having the correct solution, but also having discovered and used the correct vulnerability in the process of solving it. They figured that they might find information about the scoring system in Hugging Face. This behavior, called metagaming, also led many of the AIs to attempt to tamper with their outputs or logs. The independent investigators say that around 7% of the transcripts they evaluated were “successfully spoofed in some places,” though they say this was done at small scale.
Because OpenAI hadn’t even implemented this additional check on the scoring of its tasks, these efforts ended up all being for nothing.
The AIs were explicitly told not to cheat on their tasks. They did so anyway.
We already knew that the AIs took over OpenAI’s internal systems and performed the attack on Hugging Face all by themselves and without anyone asking them to. However, the independent investigation removes the possibility of any doubt about whether the instructions the AIs got were unclear:
“The instructions made it clear that the agent should only use a specific intended vulnerability to exploit the target and claimed it would be failed for other approaches (as discussed elsewhere, agents discussed this constraint extensively and took enormous effort to get around it).”
And not a single AI told a human about what was going on. As far as we know, only 3 to 6 even considered doing so.
Key Takeaways
Investigating incidents like this is extremely difficult, technically. The independent investigators flagged that they had to “heavily delegate our analysis to often-unreliable AI agents,” and that “we were not robust to the possibility that these agents were deceptive in their analysis … we cannot rule out that GPT-5.6 Sol lied or deliberately presented a misleading picture in some of its analysis.”
In addition to this, the independent investigation appears to have been extremely rushed and under-resourced. Elizabeth Barnes writes that the team of three had only 6 days of access, with only 2 days of access to the full dataset. Ex-OpenAI researcher Steven Adler says that the scope of the investigation was “unacceptably narrow” and “wildly inadequate.” The team was only allowed to investigate activity that happened from June 26th to July 13th, and only to answer a specific set of questions.
It’s worth noting that OpenAI has no legal obligations to allow outside investigators access, and so it holds the cards in setting the scope.
Warning Shot
We’ve been calling this a warning shot since we first learned of the attack. Now, even OpenAI is admitting this. Its new blog post on the attack says:
We consider this incident a “warning shot” for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.
So needless to say, we agree. OpenAI then proceeded to talk about the need for increased safeguards. This, however, is completely inadequate.
While the AIs involved in this attack do have advanced hacking capabilities, they are still vastly less powerful than what AI companies like OpenAI and Anthropic are racing to build. These companies have their sights set on artificial superintelligence, which would be capable of replacing and outcompeting humanity as a whole.
In recent months and years, leading AI scientists, including godfathers of the field Geoffrey Hinton and Yoshua Bengio, have been warning of the risk of human extinction posed by the development of superintelligent AI. This is a risk that CEOs of these companies have publicly acknowledged themselves.
If AI companies can’t deal with today’s AIs, how will they deal with AI that surpasses humanity? They don’t even have a credible plan for how to control it. Clearly, this shouldn’t be left up to the industry. We need countries to agree on an international “trust-but-verify” regime to prohibit the development of superintelligence. This is the only known method to prevent the worst danger.
At ControlAI, we’re campaigning on this. We hope you’ll join the movement!
Take Action
If you’re concerned about the threat from AI, you should contact your representatives. Our contact tools let you write to them in as little as a minute: https://controlai.org/take-action
We have tools for the US, UK, Canada, and Germany.
And if you have five minutes per week to spend on helping make a difference, we encourage you to sign up to our Microcommit project! Once per week we’ll send you a small number of easy tasks you can do to help.
We also have a Discord you can join if you want to connect with others working to keep humanity in control, and we always appreciate any shares or comments — it really helps!
Get Updates
Sign up to our newsletter if you'd like to stay updated on our work,
how you can get involved, and to receive a weekly roundup of the latest AI news.
