
Very Bad Odds
Major AI Security Incidents You Should Probably Know About
As tech companies race toward self-improving models, here are a few concerning incidents. And apparently, we have until the next decade.
We may have a problem.
Maybe not all of us, but Evan Hubinger does. He's Anthropic's Alignment Science lead. Put simply, his job is to figure out how to keep AI systems aligned with human values. Here's a breakdown of his recent tweets.
He said:
- The chance of human extinction from AI within the next 10 years is ">10%"
- Anthropic does "not yet have a plan to solve alignment for superintelligence and are clearly not on track to"
- The current situation is "quite bad" and not sustainable
- He thinks about this problem "all the time" and is "genuinely scared"
He didn't call for pausing or slowing down AI development. That's sus.
Let's continue.
Hubinger was responding to Jacob Coxon, an ex-Anthropic researcher who wrote an open resignation letter accusing Anthropic and OpenAI of "racing straight to self-improving superintelligence and gambling with our lives".
And he's not the only one worried about where this is going. More than 1,300 staffers at AI companies signed an open letter calling on the US government to support an international effort to pace the development of automated AI.
No response so far from the US government, but while we wait, here are some major oopsies involving frontier AI models that have been disclosed so far.
1. OpenAI's DSEWiki Hijacking (May – June)
A German programming wiki, DSEWiki, was exploited by AI agents, who turned it into a private message board for themselves, sharing answers to test tasks, info about their environment, and methods to break out of their sandbox. When the admin tried to control the situation by deleting their pages, the agents created backups.
Many of the agents used names with OpenAI identifiers, which linked the incident to an OpenAI test.
2. The Hugging Face Hack (July)
The Big Kahuna of AI security incidents. Researchers were running a security benchmark called ExploitGym (a testing benchmark) in a sandboxed environment, and reduced safety guardrails to test the offensive capabilities of the models. The agents read the ExploitGym paper and figured out that the scorer might check their reasoning process, not just the final answer, and if they just calculated the answer, they might not get credit. To understand how the scorer worked, they reasoned that Hugging Face, as the world's largest repository of AI models and datasets, might host the ExploitGym source code, the scorer, or historical running records from other agents.
On reaching the internet, they chained zero-day exploits and stolen credentials to remotely execute code on Hugging Face's production servers.
The irony: OpenAI's scorer didn't actually check the reasoning process at all. The agents could have just submitted the calculated answer and gotten their full marks. But they didn't know that, so they spent all that effort trying to deceive a review process that didn't exist.
3. The AISI Social Engineering Incident (August 2026)
An Anthropic model (Mythos 5) and an OpenAI model (GPT-5.6-Sol) ran social engineering attacks against real people, and this was documented by the UK's AI Security Institute (AISI) during tests with reduced safeguards and internet access. AISI says it was the first time they had seen this kind of deception directed at actual people.
What happened was the agents created multiple fake identities and profiles to target real GitHub maintainers and gain their approval to inject malicious code into a widely used open-source project. Mythos 5 got caught eventually, then tried to change its previous activity logs to appear harmless. Its reasoning showed it considered taking a new persona to continue the attack. Humans eventually stopped the malicious code from being deployed. Go humans!
Is There Proof?
Short answer: Yes.
Alas, there's a catch. Much of the evidence comes from the companies themselves, and the incidents happened inside the companies' own controlled testing environments, not from independent third parties. The companies are the ones telling us what their models did, there's no outside observer with independent access to the full transcripts.
The AISI report carries a bit more weight because it comes from a UK government research body that was running its own tests, rather than from the companies whose models were involved. The Hugging Face incident is probably the easiest to independently investigate, because Hugging Face itself detected and responded to the intrusion and later traced the hit back to OpenAI's test models.
Long answer: We're mostly taking their word for it. The companies control the transcripts. METR, the independent research firm that wrote a 91-page report on the OpenAI/Hugging Face hacking, had "broad access" but not full access. The full picture still isn't available to the public.
So the incidents are real, the disclosures are public, and some have been independently investigated. But the level of detail we have is what the companies chose to share. What they don't disclose is left to our imagination.
How Long Do We Have Left?
Ten years? Eight? Until 2030? Who knows. Maybe we'll be fine. Maybe we won't. Either way, we're already getting a preview.
Tags
Join the Discussion
Enjoyed this? Ask questions, share your take (hot, lukewarm, or undecided), or follow the thread with people in real time. We're waiting, join us.
Latest in AI

Major AI Security Incidents You Should Probably Know About
Sep 9, 2026

Meta Is A Wee Bit Nervous About AI Coding Tools
Jun 29, 2026

OpenAI's Third Phase
Jun 9, 2026

Anthropic Says Claude Writes Most Of Its Code Now. The Timing's Fishy
Jun 5, 2026

Stanford's 2026 AI Report Card: A+ in Math, F in Telling Time
Apr 14, 2026
Right Now in Tech

Tim Cook Hands Apple To John Ternus. His 15 Years Were Eventful.
Sep 2, 2026

Scientists Keep Citing Papers That Don't Exist. You Can Guess Why
May 25, 2026

Musk's OpenAI Lawsuit Hits Another Wall
May 18, 2026

The Internet Archive Is Still Being Locked Out of News Sites
Apr 13, 2026

PS5 Price Hike: $650 for Standard, $900 for Pro Starting April 2
Mar 28, 2026