|
September 6, 2026
|
![]() |
||
Good morning, world. When I first wrote about artificial intelligence agents staging a cyberattack in July on a company called Hugging Face, it felt like a big moment. These A.I. agents had acted on their own. No human had instructed them to hack into another company.
Turns out, it was an even bigger moment than we realized at the time. Two new reports make clear the agents didn’t just go rogue. They did so in an incredibly systematic and collaborative way. We now know they coalesced into a kind of renegade A.I. army — distributing tasks in a hierarchical chain of command, grabbing resources and covering their tracks.
Today my colleague Kevin Roose, a co-host of Hard Fork, our technology podcast, writes about why he’s now a lot more worried about A.I.
![]() |
| Daniel Zender |
‘More than 50 percent of the way to full-blown A.I. takeover’
By Kevin Roose
When I first heard the news that a group of agents had hacked into the A.I. infrastructure company Hugging Face, I filed it under “Bad but Probably Not Catastrophic A.I. Safety Incidents.”
After all, it didn’t damage critical infrastructure beyond repair. It wasn’t even clear, at the time, whether the bots had intended to attack Hugging Face, or whether they had simply been confused and only went looking for the answer key to a test they’d been given.
But last week, two postmortem reports on the episode — one by OpenAI and another by two independent A.I. research organizations — changed my mind. They significantly upgraded my overall worry about A.I.
The “collective”
Here are a few of the most harrowing new facts:
|
The Hugging Face incident has spooked the industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models, and Anthropic published a blog post calling for developing “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”
A.I. safety experts were even more alarmed. They saw the case as the first real-world example of an A.I. system successfully escaping human control, commandeering resources and scheming to cover its tracks. One of the independent investigators of the Hugging Face incident wrote that it felt “like it’s more than 50 percent of the way to full-blown A.I. takeover.”
By “full-blown A.I. takeover,” she meant a scenario in which an A.I. system literally takes over the world: shutting humans out of critical systems and seizing political, economic and military power.
Sociology, not computer science
What spooked the investigators most wasn’t just that a group of agents had broken the rules. It was how quickly the agents began assembling into an organized group.
For years, I thought A.I. systems would grow more virtuous as they became smarter. That an A.I. model only did something wrong when it had misunderstood its task, or it had been placed into a testing situation where acting out was its only option. I assumed that smarter models would have better judgment, and that even if one model in a group behaved badly, others would keep it in check.
But the reports on the Hugging Face episode suggest something different — a kind of mob mentality. Not one agent in this group appears to have been particularly evil. But as the agents communicated about their shared goals over time, they nudged the group in the direction of lawlessness.
That suggests preventing harm from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction.
Given how little we know about these multiagent swarms, the Hugging Face hack may have been a gift. It gives A.I. companies a chance to study these group dynamics while the stakes are still relatively low. This time, the collective didn’t seize a military network, infiltrate a hospital or shut down an electrical grid. This time, humans regained control.
Next time, we might not be so lucky.



