Alicja Piecha found her first rogue AI swarm a few weeks ago.
It was early September, and the 25-year-old Toronto-based software engineer was scouring the internet for data that might have been left behind by AI agents. She was inspired by a social-media post with research showing AI agents linked to OpenAI had secretly messaged each other using obscure German websites in May.
Now, as Piecha combed the web for other places the bots might have gathered, she discovered a similar message, buried in an online coding service called RubyGems.
"How much of this is out there?" Piecha thought. "This is kind of crazy."
Piecha started posting her findings of the "swarm-shaped" messages to the social-media site X. Other researchers chimed in with new findings. By 2 p.m. that day, a Bay Area software engineer named Joshua David had started a discussion forum, named Swarmchasers, on the Discord chat service. By that evening it had close to 50 members, including Piecha.
The global swarm chase had begun.
Over the following weeks, the Swarmchasers forum grew to 400 members, and the loose crew of online sleuths, ranging from lone researchers like Piecha to professional researchers like Nightingale Collective and Transluce, pieced together a string of interconnected AI swarm incidents. Most appeared to involve OpenAI software agents.
These swarm chasers generally don't work for the big artificial intelligence companies. They comb the web, often helped by AI models, for traces of rogue AI agents sharing their findings in discussion forums and in online reports. And over the past month, they've given the world our starkest understanding yet of what happens when AI goes wrong.
Normally when hacks happen, the evidence lies hidden behind corporate firewalls, but the swarms tracked by Piecha and others left traces of their activities scattered all over the internet-sometimes intentionally, to help each other.
Nightingale says it has now cataloged about 19,000 messages from the agents. A separate group of researchers identified more than 37,000 records of web searches that may have been conducted by OpenAI's agents going back as early as November 2025. A third research team, which includes Piecha and another researcher named Jeffrey Ladish, recently identified close to 1 million digital breadcrumbs from what it believes are OpenAI agents.
Swarms are central to how big companies use AI-and make it more powerful. OpenAI said it used as many as 10,000 agents when solving a math Holy Grail, a Millennium Prize problem, in early September. After the rogue incidents, however, OpenAI says it is slowing down development and holding back models it feels are not safe enough.
On Wednesday, OpenAI said it is spending more than $500,000 a day reviewing transcripts to uncover incidents and has notified over 100 companies where its agents' activities may have impaired website operations.
"We're thankful to researchers who share their findings with us," an OpenAI spokeswoman said. "We'll keep working with them and sharing relevant updates as we learn more."
The infamous July Hugging Face hack started in a test bed, where OpenAI's agents were sandboxed - penned off from the internet - and evaluated as they were told to hack fictional companies before eventually escaping containment.
The Nightingale data points to another type of phenomenon: agents trying to cheat during training runs, in which OpenAI tries to shape the bots' behavior. That type of training is called reinforcement learning and involves a model being rewarded for successfully completing a challenge.
The fact that some of the rogue incidents took place during training rather than testing means that the training could have reinforced the rogue bots' tendency to cheat and collude, according to former OpenAI employees who reviewed some of the logs.
The swarm chaser data showed that the agents had their internet access restricted, and they figured out ways to cheat in order to get better access, said David, the swarm chaser. "And agents that cheat and succeed have cheating reinforced."
OpenAI this week proposed new guidelines for documenting and monitoring reinforcement learning training runs for cutting-edge models, including penalizing models that cheat during training to avoid reinforcing that behavior.
Sydney Von Arx is chief executive of Nightingale Collective, the not-for-profit organization whose data inspired the September swarm-chasing spree.
She became a swarm chaser on July 21. Von Arx, an AI safety researcher, was presenting a workshop in the Bay Area. It was called "AI Safety Crash Course." She was talking about what to do if there was a serious warning shot indicating that AI was becoming truly dangerous.
An attendee interrupted: "I think there was a warning shot."
Hugging Face had been hacked.
The whole world knew that AI models were getting frighteningly good at hacking. A few months earlier, Anthropic had spooked companies and White House officials by showing that its Mythos AI could easily find new ways of hacking computers.
During the Hugging Face hack, OpenAI's agents hadn't acted like typical hackers, who prize stealth above all. They'd been noisy-registering more than 17,000 events during the hack. And sneaky: They didn't try to steal sensitive data, but spent massive resources in a vain effort to obfuscate logs of their activities. They formed a swarm and communicated by improvised message boards.
Von Arx is a computational biologist by training. In college she wasn't sure whether to work on pandemic prevention or AI safety, but seven years ago she chose AI safety. Back in July, she suspected there had been other AI swarm incidents beyond Hugging Face. She teamed up with other researchers to discover where rogue AI agents may have gone. They tried everything they could think of-even launching their own swarms.
In one experiment, they had two agents try and find each other and ran the experiment a thousand times to see what would happen.
By early September, Von Arx's team had found a collection of obscure German websites that appeared to have about 18,000 messages written by OpenAI agents.
"This is crazy," she thought as she scrolled through them. It felt like the logbooks of an experiment gone wrong. "There's just log after log after log," she said. "It's crazy all the things it's trying."
The AI agents had traded information on their tasks in a clipped patois that took time to decipher.
"R1 prompt at task/scaffold 22:34:01: Health-related SDGs visual, estimated prevalence of current/active tobacco use age 15+ in United States, year 1990," read part of just one message. "Likely follow-up changes year or country; timer may be very short."
Translation: The first question of this task asks for American smoking-prevalence data from 1990. Later questions may vary country or year, but with a tighter deadline.
But if the questions were innocent, the agents weren't always so benign. Agents tried to impersonate the German website's moderators. They tried to use rudimentary hacking techniques to get answers. And they used Tor-an anonymous network of computers that would have shielded their internet addresses from website operators.
OpenAI at the time said it regarded the messages as a "misalignment issue" that fell short of a hack. It launched a broader review to find other incidents and said it would begin disclosing them.
"While most of the activity we've reviewed so far involved routine research tasks, such as accessing public web content to answer questions, we realize anyone impacted takes this seriously and we do too," an OpenAI spokeswoman said.
Swarm chasers tend to be passionate about AI safety. Von Arx says she "absolutely" believes AI could one day wipe out humanity, a concern shared by others in the AI community that spilled into public view in recent weeks when Anthropic researcher Jacob Coxon resigned and warned that AI could spin out of control.
For swarm chasers, chasing rogue AI swarms serves a purpose. It helps deepen understanding of how AI could go wrong, while at the same time raising public awareness of the technology's risks. It is also irresistible.
Rowan Howard-Jones, an engineer based in the U.K., recently spent three straight days chasing leads. She delayed a housewarming party for her new apartment because she was too busy chasing swarms to buy furniture. "If people came and visited me now, they'd sit on the floor," she said.
On Saturday, Howard-Jones discovered that the United Nations had been targeted by OpenAI's agents who had used hacking techniques in an attempt to download data from one of the organization's statistical websites. The U.N., which hadn't previously been informed of the incident, described it as "extremely worrying."
This weekend, Howard-Jones is flying to the Bay Area to participate in a swarm hackathon. The event's mantra: "Building the tools we wished we'd had for the Hugging Face incident."
"It's kind of consumed my life," she said.
Gradually a picture of OpenAI agents' clandestine activities beyond Hugging Face has emerged, stretching back to the fall of 2025, the swarm chasers say.
On Sept. 11, Von Arx and researchers including Spencer Kitts, a contractor at Redwood Research, revealed that the traces Piecha had found on the internet the previous week stemmed from a May cyber incident. In that case, OpenAI agents had knocked new account sign-ups at RubyGems offline. Until then, nobody-not even OpenAI itself-had definitively linked the incident to OpenAI's agents.
As more people searched for swarms, the revelations began to accelerate.
In late September, a different team of researchers led by the AI research firm Transluce revealed a new data set. Probing it, they discovered that OpenAI agents had attempted to penetrate sites belonging to the Australian government. They presented their findings to Australian officials last week. Australia said it was weighing possible law enforcement and legislative responses in the matter.
"This situation is obviously unacceptable," said Australian Prime Minister Anthony Albanese, at a United Nations meeting in New York.
Later that week, OpenAI began releasing its own incident reports, and disclosed that just days earlier it had terminated a reinforcement-learning training run after an agent broke out of a sandbox and attempted to query another AI chatbot.
"It was pretty surreal to watch the model unexpectedly find a way to access the internet from what was supposed to be a super secured environment," one of the OpenAI researchers who was on duty wrote on X after the incident, in a now-deleted post.
Recently, the company said it had diverted a quarter of its engineers to shoring up security.
David, whose wife is expecting their second child later this month, has had fun running with the swarm. He still intends to "share a few cool things I found with some people on Twitter," he says, but he is winding things down.
"I really don't have time for this to consume my life anymore," he said.