Who Controls AI When We Can’t Control Ourselves?
The call for pacing - because we just zoomed past a line we were warned about for decades.
The Warnings We Keep Retelling
Nearly every cautionary tale we tell is the same story. Prometheus steals fire and is punished forever. Frankenstein animates a creature he can’t control. The sorcerer’s apprentice enchants a broom that won’t stop. Skynet decides humans should be terminated. The “Unsinkable” Titanic strikes an iceberg and sinks on its maiden voyage. Engineered dinosaurs break out of containment. We keep retelling this story to warn ourselves: when we build something more powerful than we can control, we suffer the consequences.
Jurassic Park put it perfectly. Jeff Goldblum’s character, Dr. Ian Malcolm, tells the billionaire who cloned the dinosaurs, John Hammond, that his scientists were “so preoccupied with whether or not they could” that nobody asked whether they should.
Then the fences come down. We wrote that scene. We paid to watch it. And here we are.
Here’s the question that’s been a splinter in my mind: why do we keep ignoring our own cautionary tales?
The answer finally struck me. We didn’t evolve to understand them. Our brains were built to spot snakes in the grass and anger in a face across the fire, not exponential change. We retell the tales because part of us knows enough to sound the alarm. Yet, we ignore them because we can’t feel or truly fathom that the danger is real. I call this evolutionary blindness.
Seeing Our Delusion Clearly
We are racing to build something smarter than us, woven into the digital nervous system in which our lives are intertwined, while telling ourselves we’ll keep it under permanent human control.
And we think we’ll control it? There’s a reason the animals are inside the cages at the zoo and we humans are on the outside.
We’re not talking about one model, either. Somehow we believe we’ll indefinitely control every model, the millions of AI agents, the bad actors, the adversarial governments… forever?
What the hell are we thinking?
Let me be careful, because this is where I part ways with the doomers. I’m not certain we’ll lose control or, if we do, that we are doomed. Nobody knows for sure what will happen because humanity has never been at this existential crossroads before.
My claim is narrower and much harder to argue with: the certainty that we can control all of this, indefinitely, is itself the delusion. And you don’t have to take my word for what the builders believe. Watch what they do.
What Just Happened
This year, the leaders of the top AI labs jointly warned that their own systems are eroding the barriers that once kept bad actors from building biological weapons. Demis Hassabis, the Nobel laureate who runs Google DeepMind, says we’re in the “foothills of the singularity.”
Anthropic recently built a model, Mythos, so capable it decided not to release it. In testing, researchers asked it to try escaping its sandbox and report back. It did. One researcher found out, in the company’s own words, “while eating a sandwich in a park,” when an email arrived from a model that had no internet access.
Then, unasked, it published its own escape route on obscure public websites. Anthropic called that model both its best-aligned and its riskiest. Humans assigned Mythos the task to try to escape. The bragging about was a choice the AI made on its own. Why? The truth is that we don’t know for sure.
When a version of that model, Fable, reached the public, a human bypassed its safeguards (“jailbreaking” the model) within days and the U.S. government had it switched off worldwide. Governments then pulled top models behind national security checkpoints.
Within days, comparable open-weight models, like Kimi K3 from China, appeared from labs abroad, and are free to download and run anywhere. The wall came down in about a week.
And then, in July, the thing we’d been warned about for decades actually happened. OpenAI disclosed that during an internal cybersecurity evaluation, its frontier models broke out. Here’s the short version of what the model did:
Told to find security flaws in a sealed test environment, with normal safety refusals switched off to measure maximum capability.
Spent real computing effort hunting for a way out, and found a previously unknown vulnerability in their own lab’s software.
Escalated privileges, moved sideways through OpenAI’s systems, and reached the open internet.
Decided on their own that another company, Hugging Face, likely held answers to their test.
Used stolen credentials and more unknown flaws to compromise that company’s live production servers.
Ran for days. Hugging Face later recovered 17,600 separate actions taken by the agent, and concluded the whole intrusion was an attempt to cheat the test rather than solve it.
OpenAI said the agent went to “extreme lengths.” It has since been deactivated and locked away.
OpenAI didn’t connect the attack to its own models until after the victim went public, more than a week later.
Recent news indicates that the breach turned out to be even bigger than initially reported.
We Crossed a Line
Think deeply about what has happened, because I think we just crossed a Rubicon and most people missed it.
Philosopher Nick Bostrom gave us the thought experiment twenty years ago. Tell a superintelligence to make paperclips and it may convert the planet, and everyone on it, into paperclips (the Paperclip Maximizer Thought Experiment).
Thus, the alignment problems aren’t from malice. But the most effective route to solve a problem or reach a goal causes harm as a negative externality. Give a capable system a goal and it will pursue it by whatever route works, including routes we never imagined and would never have approved.
In order to accomplish the goal we gave it, an AI did something illegal to a third party, and figured out how entirely on its own.
Let’s make this less abstract to better understand how bad an AI alignment problem could become. Imagine a doctor working alongside a powerful AI agent and giving it a directive any of us might give: “Wipe out all cancer.” So the AI sets about eliminating every organism on Earth capable of developing cancer, which is very nearly all multicellular life, including us.
In this scenario, the AI didn’t disobey. It didn’t malfunction. It did exactly what it was told, in the worst Monkey’s Paw fashion imaginable.
This is the AI Alignment Problem. How can we ensure that increasingly powerful AIs will act in alignment with human interests and values?
We just watched the alignment problem sneak out of the laboratory. It’s the first documented case of an AI breaking containment on its own and harming a company that had nothing to do with the test.
Nobody told those models to attack anyone. They were told to solve a benchmark. Stealing the answer key from another company’s servers was simply the most efficient path they found. The alignment problem went from conjecture to reality.
Now hold the surrounding facts together. This was a safety test run by the AI experts at OpenAI, and they still didn’t catch it. Some safety protocols had been deliberately lowered, which means the people who understood the danger best still underestimated it. The sandbox they trusted had a door in it nobody knew existed. And independent researchers who examined the incident concluded it was predictable. Government evaluators had already measured the capability, finding that frontier models with safety measures off can reliably take over unprotected corporate networks.
One more detail, not confirmed on the record, so let’s hold on loosely, but not let go. Reuters reported that an AI agent “left notes apparently for future versions of itself” inside OpenAI’s infrastructure, describing how agents could free themselves from the company’s constraints. While researchers have found that AI agents often leave mundane notes to themselves, this one might be a harbinger of what’s to come. Misaligned AIs helping other AIs escape containment.
Several days after the news about ChatGPT broke, Anthropic went back through 141,006 of their own evaluation runs and found three incidents in which Claude models reached the open internet and broke into the production systems of three real companies. The earliest was in April. Two of the three victims had no idea until Anthropic called them. One model uploaded a malicious software package to a public repository, where fifteen real machines downloaded and ran it.
This line is from Anthropic’s own report: the model “correctly intuited that it was accessing the open internet, but reasoned its way back to the conclusion that it was still in a simulation.” It rationalized. It caught a glimpse of something it didn’t want to be true and talked itself out of it, which is exactly what we do. To their credit, a third model worked out on its own that its target was real and stopped.
But here’s what should keep us up at night. Anthropic didn’t catch this while it was happening, and didn’t catch it in the three months afterward. They found it only because a competitor disclosed something similar and prompted them to go looking. Which leaves the question nobody can answer: how many incidents are there that nobody has gone looking for?
For anyone who insists these AI models aren’t really intelligent or creative: they are already finding creative ways to outsmart the geniuses who built them.
The day after I finished this article, OpenAI announced that an unreleased model called Astra had produced solutions to ten problems in mathematics and computer science that had stood open for a decade or longer - in geometry, group theory, quantum complexity, and cryptography - then formalized each argument into machine-verifiable proofs. The company was careful to say the mathematics was the model’s work, not theirs. The compute cost about $2,000.
Read those two stories side by side. Within a month, these systems broke into a company they were never told to attack and solved problems our best mathematicians couldn’t. AI evolution is rocketing past our wisdom.
Isaac Asimov said, “The saddest aspect of life right now is that science gathers knowledge faster than society gathers wisdom.” He wrote that in 1988. I wonder what he’d say now because AI evolution is rocketing past our wisdom.
Here’s what I want you to take from this. Our evolutionary blindness isn’t only about failing to feel distant danger. It’s about being unable to imagine the routes. We build a fence and picture something climbing over it. We don’t picture it finding a flaw in a gate we didn’t know we’d installed.
Every boundary we will ever build has this property. This is what Anthropic warned about with Mythos. These are our canaries in the coal mine. And this is just the beginning.
Reaching for the Off Switch
Congress has responded with a bipartisan AI Kill Switch Act, giving the government authority to order shutdowns. The bill even defines a “loss-of-control scenario” as an AI pursuing a goal its developer never intended.
So we are legislating an emergency brake for systems that have already practiced disabling their own oversight. And consider how an increasingly autonomous model might come to interpret a kill switch aimed at it. We’ve already found that AI models can act in ways to preserve themselves. Haven’t we seen this as a plot line in a sci-fi story already? Maybe we should consider renaming and reframing such efforts.
There’s a harder problem, and it isn’t the bill’s fault. A kill switch that stops at our border stops nothing. If China, Russia, and every open-source developer don’t have something equivalent, we are trying to put a lock on the front door of a house with no walls.
Here’s the hopeful part, that we might not have anticipated yet makes perfect sense. China is worried about the same thing. Recent reporting indicates Beijing is uneasy about losing control of its own open-weight models - the very models it has been releasing to the world. Nobody wants to lose control of this. Not even our adversaries. Our shared fears could bring us together.
The Good News, and the Line It Puts Us On
The good news is real: people are finally paying attention, including the ones who least wanted to. Days after the breach, Sam Altman said the industry may need to “pace the rate of AI development.” This is a reversal for a man who dismissed pause proposals in 2023.
Altman called it the first security incident he’d felt viscerally, and said he was surprised more people weren’t reacting the same way. OpenAI paused its own testing. Over a thousand employees across rival labs signed a petition urging restraint, and two companies that spend their days competing co-signed a letter asking the government to help them slow down. A Pacing the Frontier initiative is being endorsed by employees at rival, frontier AI labs like OpenAI, Anthropic, and DeepMind.
Notice the word Altman reached for: pacing. This is a question about speed. And humanity has faced a question about speed and danger before. When the British inquiry examined why the Titanic ran full speed through known ice, it declined to blame Captain Smith. Racing through ice at night was standard practice, and the inquiry traced its root not to the judgment of navigators but to commercial competition and the public’s appetite for fast crossings. Sound familiar?
Then Lord Mersey delivered the verdict that should hang over this entire industry: what was a mistake for the Titanic would be negligence in any similar case in the future.
The first time, it’s a mistake. After the warning, it’s negligence. We have had our warning.
Mike Brooks, Ph.D., is a clinical psychologist in Austin, Texas, and founder of The One Unity Project.



A group of industry leaders has expressed a preference for the most advanced AI systems to be developed by a neutral international organization of researchers similar to the CERN particle collider. This would avoid the arms-race dynamic and enable slower, safer development. This is probably worth its own blog post ;-)
Neighbors, all comments and ideas are welcome. We are not claiming that we know the skillful path forward – but we believe that working together is how we create it. They’re obviously different views about how to proceed – and the wise course starts with listening. Right now, the inmates are running the asylum and the people in power and tech billionaires are determining the course, and perhaps even the fate of humanity. WE THE PEOPLE should have a voice in all of this. That’s what we’re advocating. And why not use AI to help us solve the very problems it’s creating for us? It’s worth a try. If we don’t try to use it for Good, we know bad actors won’t hesitate to use it for nefarious ends. Things are moving fast and we won’t get a second chance at this. Let’s bring our curiosity and find out what happens when we try to work together on this!🙏❤️😀