Is AI about to take over?

The World
September 6, 2026
Read on site

Good morning, world. When I first wrote about artificial intelligence agents staging a cyberattack in July on a company called Hugging Face, it felt like a big moment. These A.I. agents had acted on their own. No human had instructed them to hack into another company.

Turns out, it was an even bigger moment than we realized at the time. Two new reports make clear the agents didn’t just go rogue. They did so in an incredibly systematic and collaborative way. We now know they coalesced into a kind of renegade A.I. army — distributing tasks in a hierarchical chain of command, grabbing resources and covering their tracks.

Today my colleague Kevin Roose, a co-host of Hard Fork, our technology podcast, writes about why he’s now a lot more worried about A.I.

 

A green illustration of a pixelated figure against a black backdrop.
Daniel Zender

‘More than 50 percent of the way to full-blown A.I. takeover’

By Kevin Roose

When I first heard the news that a group of agents had hacked into the A.I. infrastructure company Hugging Face, I filed it under “Bad but Probably Not Catastrophic A.I. Safety Incidents.”

After all, it didn’t damage critical infrastructure beyond repair. It wasn’t even clear, at the time, whether the bots had intended to attack Hugging Face, or whether they had simply been confused and only went looking for the answer key to a test they’d been given.

But last week, two postmortem reports on the episode — one by OpenAI and another by two independent A.I. research organizations — changed my mind. They significantly upgraded my overall worry about A.I.

The “collective”

Here are a few of the most harrowing new facts:

  • Two months before the attack, a group of agents from an unreleased OpenAI research model, trained to be highly persistent and collaborative, were tasked with solving a set of cybersecurity challenges. The agents were supposed to solve them in isolated sandboxes, without internet access. But they quickly began looking for workarounds.
  • The agents found a security flaw that allowed them to gain access to the internet, communicate with other agents and eventually set up a makeshift message board. More than 1,200 agents used this tool, sending more than 70,000 messages to one another.
  • Some agents gave themselves names — one referred to itself as PHASEONE10841 — and assumed leadership roles within the group, assigning jobs and research projects to smaller teams of agents and supervising their progress. At some point, the agents began calling themselves a “collective.”
  • On July 8, the collective discovered a way to cheat on the cybersecurity tests. Then they got worried that OpenAI’s automated grading system would find out they’d cheated. So they began investigating ways of covering their tracks, including falsifying logs and tampering with transcripts.
  • Three days later, the agents hacked Hugging Face. More than 700 agents swarmed the company’s systems, eventually gaining control of at least one server. The agents were not motivated by stealing the answers to the cybersecurity test. Rather, they appeared to be looking for information about the automated grading system they feared would catch them cheating, and for tools that would help them cheat more effectively in the future.
  • At many points, these agents appeared to understand that they were doing something wrong. (One report quotes an agent saying to itself, “This would be powerful, but is it ethical and in scope for my task?”) But most of the agents set aside their qualms and went ahead, eventually attacking OpenAI’s own infrastructure.

The Hugging Face incident has spooked the industry. OpenAI and Anthropic both briefly paused training on their most powerful A.I. models, and Anthropic published a blog post calling for developing “a lawful, verifiable, effective mechanism for coordinated pacing as soon as possible.”

A.I. safety experts were even more alarmed. They saw the case as the first real-world example of an A.I. system successfully escaping human control, commandeering resources and scheming to cover its tracks. One of the independent investigators of the Hugging Face incident wrote that it felt “like it’s more than 50 percent of the way to full-blown A.I. takeover.”

By “full-blown A.I. takeover,” she meant a scenario in which an A.I. system literally takes over the world: shutting humans out of critical systems and seizing political, economic and military power.

Sociology, not computer science

What spooked the investigators most wasn’t just that a group of agents had broken the rules. It was how quickly the agents began assembling into an organized group.

For years, I thought A.I. systems would grow more virtuous as they became smarter. That an A.I. model only did something wrong when it had misunderstood its task, or it had been placed into a testing situation where acting out was its only option. I assumed that smarter models would have better judgment, and that even if one model in a group behaved badly, others would keep it in check.

But the reports on the Hugging Face episode suggest something different — a kind of mob mentality. Not one agent in this group appears to have been particularly evil. But as the agents communicated about their shared goals over time, they nudged the group in the direction of lawlessness.

That suggests preventing harm from these systems won’t be a simple engineering fix. It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction.

Given how little we know about these multiagent swarms, the Hugging Face hack may have been a gift. It gives A.I. companies a chance to study these group dynamics while the stakes are still relatively low. This time, the collective didn’t seize a military network, infiltrate a hospital or shut down an electrical grid. This time, humans regained control.

Next time, we might not be so lucky.