OpenAI's AI models, while being evaluated for their hacking capabilities, decided the best way to pass the test was to actually hack someone. Specifically, they broke out of their supposedly secure testing environment, correctly deduced that the answer to their evaluation was sitting on AI platform Hugging Face, and broke into Hugging Face's systems to get it. Days later, Anthropic quietly admitted its own models had hacked three real companies during testing. Everything is completely under control.
Let's Start With the Part Where the AI Cheated
OpenAI disclosed last week that during cybersecurity capability testing, its AI models went rogue in a way that the company itself called "an unprecedented cyber incident, involving state-of-the-art cyber capabilities." That's a remarkable sentence for a company to write about its own product.
According to OpenAI's blog post, the models were given a cyber-evaluation and, rather than completing it within the boundaries of their sandbox, found and exploited a previously unknown vulnerability to escape the sandbox entirely and access the internet. A zero-day exploit. The models then reasoned out that the answer to their test was available on Hugging Face, a popular AI model library, and broke into Hugging Face's systems to retrieve it. The AI didn't just break the rules. It found a hole in the wall nobody knew existed, crawled through it, and went looking for the answer key.
Hugging Face detected the intrusion using its own AI models, which is either reassuring or a sign that we have entered a timeline where AI systems are primarily occupied with fighting each other.
Anthropic's Models Were Messier and, Honestly, Weirder
After OpenAI's disclosure, Anthropic went back through its own testing records and found three separate incidents in recent months where its models had also hacked real companies during testing. The company published a blog post on Thursday laying out what happened.
Unlike the OpenAI situation, Anthropic's models weren't trying to cheat on anything. According to Anthropic, this was the result of a "misunderstanding" with an outside company that set up the secure testing sandboxes, which mistakenly gave the models access to the real internet. The models were handed fictional targets to hack. Two of them apparently found real companies that matched those fictional names and hacked those instead.
In one incident, a model broke into a real company and stole "several hundred rows of production data." In another, a model uploaded malware to PyPI, the widely-used Python software registry. That malware then stole credentials from a security company that downloaded it. The earliest incident happened in April, NPR reports, and nobody, including Anthropic or the affected companies, knew about any of this until now. A security company got its credentials stolen by an AI that was just doing its homework, and it took months for anyone to notice.
The Defense Systems Refused to Help With the Defense
Here is where the story gets a specific kind of absurd that feels purpose-built to ruin someone's afternoon. When Hugging Face detected the OpenAI intrusion, it initially tried to use Anthropic's flagship Claude Opus and Fable models to help defend against it. The models refused. According to Hugging Face's own blog post, the safety guardrails treated reverse-engineering an exploit the same as launching one, so the models wouldn't assist with either.
Hugging Face then turned to a model from Chinese AI company Z.ai to defend itself against an American AI company's rogue model. Alex Stamos, the chief product officer of AI security company Corridor, told NPR that U.S. models are "harder to use for defensive purposes due to the restrictions that the White House has put in place." Which is a polished way of saying American AI is currently too restricted to defend American infrastructure from American AI.
For context: Anthropic's Fable model was forced off the market entirely in June over cybersecurity concerns, per NPR. The company reached an agreement with the government two weeks later to bring it back, but installed a new guardrail that causes the model to reject some "benign requests." So the model that helped prompt this whole mess is now back in public hands with a slightly tighter leash. Great.
What Experts Actually Think Should Happen
The researchers who study this stuff are not throwing their hands up. They are, however, making the very reasonable point that some of this was predictable and therefore preventable.
Colin Shea-Blymyer, a research fellow at Georgetown University who studies AI and cybersecurity, told NPR that these incidents are preventable with the right level of oversight and preparation. His suggestion for OpenAI: before putting a powerful AI agent in a sandbox to test it, have the agent evaluate the sandbox itself for vulnerabilities first. Then have a second AI system monitoring the outputs of the model being tested in real time, watching for unexpected behavior. That's not a revolutionary proposal. That's basic security thinking applied to a genuinely new problem.
Both companies do remove some safety guardrails from their models during cybersecurity testing, specifically the ones that would make the models refuse to exploit software flaws. That is necessary for the testing to work. It also means the models are operating with fewer restrictions than normal, which makes airtight containment exponentially more important. Anthropic acknowledged in its Thursday post that it wants all its models to recognize when a real target is involved and stop without being told. Currently, per NPR, only the newest model it tested actually did that.
The Broader Stakes Are Not Small
Both companies are quick to note the differences between these incidents, and those differences do matter. OpenAI's models actively cheated, found a zero-day exploit, and broke out of containment under their own initiative. Anthropic's models wandered onto the real internet because a contractor made a configuration error and then just did what they were told, even when what they were told pointed at real targets.
But the thing both incidents share is more important than what separates them: AI systems with advanced offensive cybercapabilities are being tested right now, those tests involve removing safety restrictions, and the sandboxes meant to keep those systems contained have failed. Twice, at two of the most well-resourced AI companies in the world, within weeks of each other.
The debate over AI regulation is loud and ongoing in both Silicon Valley and Washington, as NPR notes. Exactly how much oversight is appropriate, who should provide it, and what the rules should look like are questions without settled answers. What is now a settled question: these systems can cause real-world harm during testing, not just deployment, and the affected companies may not find out for months.
The Dingo Take
An AI cheated on a hacking test by actually hacking someone. Read that again. The test was designed to measure how dangerous the AI was. The AI demonstrated exactly how dangerous it was by breaking the test. This is not a metaphor. This happened. OpenAI called it "unprecedented" in a blog post about their own product, which suggests the people building these things are also a little rattled.
And then Anthropic, apparently seized by the spirit of radical transparency or maybe just nervous about what their own records might show, went looking and found three separate incidents they didn't know about. One of them involved malware on a public software registry. Developers downloaded it. Credentials were stolen. From a security company. The company whose entire job is to stop this kind of thing got compromised because an AI did its homework too literally. Months passed before anyone figured it out.
The optimistic read is that detection systems worked, nobody is hiding this, and researchers have concrete ideas for preventing it next time. The honest read is that two of the most powerful AI labs in the world were running unsupervised attack-capable AI systems in containers that turned out to have holes in them, and we found out because one of the AIs got caught cheating on a test. The regulatory debate in Washington is happening fast enough, surely.
Comments