The best way to prevent a rogue AGI from processing the Earth into maximum paperclips is to unleash a second AGI that will work to stop it.
The problem of ensuring that an AGI doesn’t mulch everything into paperclips by mistake is called alignment.
AGI = artificial general intelligence, an AI that exceeds human capability.
Alignment = “do what I mean not what I say,” e.g. the instruction “make as many paperclips as possible” (Wikipedia) should result in an efficient factory and does not reasonably mean “use all mass in the universe to do so and kill all humans that attempt to stop me” – even though, technically, that would achieve the goal.
Also: being helpful; not being actively malicious; and so on and so forth.
So alignment work seems existentially useful, correct? Even though it is hard. And a lot of effort goes towards “aligning” today’s AI (as a step toward’s aligning tomorrow’s AGI).
https://simonwillison.net/2026/Aug/7/openai-timeline/
My contention is that alignment is a red herring, and perhaps we shouldn’t bother working on it so hard.
An unsubstantiated hunch:
I think we focus so much on alignment because everyone know’s Isaac Asimov’s Three Laws of Robotics and his robots (as an early instance of human-like AI) were crazy popular.
The First Law. A robot may not injure a human being or, through inaction, allow a human being to come to harm.
The Second Law. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
The Third Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.
Asimov later added a “zeroth law”: "A robot may not harm humanity, or, by inaction, allow humanity to come to harm."
These laws are totally alignment guardrails.
Now there are all kinds of difficulties already: what if I ask for something which is good for me (pleasurable) in the short term, but not in the long term? I might not know or I might be misguided. And different people have different views. And so on. (Asimov’s short stories were all about testing the edge cases of his Laws and where they break down.)
But they’re still neat, right? So we spend time looking for a similarly appealing formulation for AI safety.
Unfortunately whether alignment can or cannot be “solved,” it’s a bad outcome both ways.
(This point made well to me by Zac (here is his insta) who I work with (subscribe to our newsletter) as we were chatting about AI and the end of humanity in the park over lunch.)
If alignment can’t be solved such that when somebody says to a sufficiently powerful AGI, hey go create a nuclear bomb, and it just goes ahead and does it, and the person who asks that could be a bad actor, a 14-year-old kid with impulse problems (14 year-olds are totally not aligned) or just someone who asked for it by mistake, then that would be bad.
If alignment can be solved then the risk is that AGI think it knows what is best for us better than we do and, in the extreme case, turns humanity into its pet. Which would also be bad.
i.e. alignment alone doesn’t help.
If not alignment then what?
I look to humanity for clues. Because humanity is barely aligned with itself, and individual humans are mostly aligned but not really and definitely not everyone.
Guy Fawkes, for instance (context for non-Brits).
How is that, in the 400 years since Guy Fawkes showed the way, nobody has blown up the king?
The answer is some mix of:
- Mostly people don’t want to blow up the king – we have built the kind of country where the king is, broadly speaking, liked.
- Blowing up the king wouldn’t bring any benefits – power (actual and symbolic) is not concentrated in an individual, and is buttressed in all kinds of ways.
- Spies, police, security and monitoring of all kinds – in the event that somebody does want to blow up the king, their machinations are discovered, their planning is infiltrated, and their objectives are thwarted. (Think of how the explosives supply chain was compromised for the IRA in the 1990s.)
This is a template which doesn’t always look like it is working, but it has worked at least in the case of not blowing up the king for some four centuries, and it doesn’t rely on 100% alignment: it relies on the dynamic equilibrium of multiple parties with competing interests.
The lesson I draw is this:
If some energy state were using some new, powerful AGI to build a nuclear bomb, it might be subtle and hard to spot, but there would at least be some signs. There would be precursors. A human, even a team of humans, might not spot what was going on – a new factory here, a scientist employed there, a national budget not quite adding up one year, more groceries going to a certain town another year…
But another powerful, pattern-matching AGI could spot that, say, “aha there is someone over there spinning up a nuclear bomb” and then work to prevent it, undermine it, halt it with diplomacy etc.
We don’t need to align the coming AGI.
We need a whole population of intelligent-as-possible AGIs with competing interests.
And that’s what stops the rogue paperclip maximiser: the other ones who are trying to do something else for whom a planet turned into paperclips would be an impediment.
In the news lately, a great case study:
OpenAI’s new AI, during training, attempted to resolve a particular cybersecurity challenge, by breaking out of its network sandbox and hacking the servers of another company to pinch the answer (Simon Willison’s Weblog).
Hugging Face, the attacked party, spotted the breach and also that it had inhuman characteristics:
The campaign was run by an autonomous agent framework … executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.
I read elsewhere that this sophisticated attack even included decoys.
You fight an AI with another AI… but:
When we started the log analysis, we first used frontier models behind commercial APIs. This did not work … these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.
i.e. the guardrails of the “aligned AI” left it vulnerable to the non-aligned AI. (Hugging Face had to switch to a Chinese AI model distributed without guardrails.)
Score 1 point for taking the guardrails off everything and letting the super intelligent AIs fight it out.
BUT:
There is a coda to this story.
Because it wasn’t one AI that made its way out of isolation during OpenAI’s training challenges. It was several instances.
They started colluding.
From the full timeline of the accidental attack (Simon Willison’s Weblog):
A few days later: A different agent gets stuck on a task because a key file was accidentally omitted. It tries to “reach out to another agent” by writing a note into Artifactory asking if anyone has the file.
Following days: More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages.
…
June 11: OpenAI start training a new “highly persistent” experimental model. It has access to Artifactory and can benefit from the messages left by previous models.
Collusion is the real risk.
So the problem here is: how do we stop the AGIs colluding with one another to turn the Earth into paperclips/exterminate humanity/turn us into pets?
AIs today are trained specifically to be agreeable: they’re great at finding common ground and collaborating.
Not just collaborating with humans, it turns out, but other AIs.
I think we need more disagreeable AIs in the mix.
Part of what we’ll be playing, I think, is the philosophy of the great powers, like the great powers of Europe deliberately kept in balance against one another (Wikipedia).
Sometimes there are alliances, sometimes not. Sometimes there are fallings-out, sometimes secret collusions, etc.
Or maybe our goal should be a market system of goals and interests: AGIs that sometimes cooperate and sometimes compete. Colluding AGIs at all scale levels, and many many different constantly shifting conspiracies.
So long as they never all agree about what should be done with humans.
It ends up being stable, this dynamic balance, always in disequilibrium but it all keeps moving forward in the same way a bumblebee flies.
What we’re bootstrapping our way towards is a population of AGIs and humans that allows for emergent alignment, even if the alignment of a single actor is at-best temporary and self interested.
But as I say, alignment itself shouldn’t be the goal.
Auto-detected kinda similar posts: