Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
158941 stories
·
33 followers

AI alignment is a red herring

1 Share

The best way to prevent a rogue AGI from processing the Earth into maximum paperclips is to unleash a second AGI that will work to stop it.

The problem of ensuring that an AGI doesn’t mulch everything into paperclips by mistake is called alignment.

AGI = artificial general intelligence, an AI that exceeds human capability.

Alignment = “do what I mean not what I say,” e.g. the instruction “make as many paperclips as possible(Wikipedia) should result in an efficient factory and does not reasonably mean “use all mass in the universe to do so and kill all humans that attempt to stop me” – even though, technically, that would achieve the goal.

Also: being helpful; not being actively malicious; and so on and so forth.

So alignment work seems existentially useful, correct? Even though it is hard. And a lot of effort goes towards “aligning” today’s AI (as a step toward’s aligning tomorrow’s AGI).

https://simonwillison.net/2026/Aug/7/openai-timeline/


My contention is that alignment is a red herring, and perhaps we shouldn’t bother working on it so hard.


An unsubstantiated hunch:

I think we focus so much on alignment because everyone know’s Isaac Asimov’s Three Laws of Robotics and his robots (as an early instance of human-like AI) were crazy popular.

The First Law. A robot may not injure a human being or, through inaction, allow a human being to come to harm.
The Second Law. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
The Third Law. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.

Asimov later added a “zeroth law”: "A robot may not harm humanity, or, by inaction, allow humanity to come to harm."

These laws are totally alignment guardrails.

Now there are all kinds of difficulties already: what if I ask for something which is good for me (pleasurable) in the short term, but not in the long term? I might not know or I might be misguided. And different people have different views. And so on. (Asimov’s short stories were all about testing the edge cases of his Laws and where they break down.)

But they’re still neat, right? So we spend time looking for a similarly appealing formulation for AI safety.


Unfortunately whether alignment can or cannot be “solved,” it’s a bad outcome both ways.

(This point made well to me by Zac (here is his insta) who I work with (subscribe to our newsletter) as we were chatting about AI and the end of humanity in the park over lunch.)

If alignment can’t be solved such that when somebody says to a sufficiently powerful AGI, hey go create a nuclear bomb, and it just goes ahead and does it, and the person who asks that could be a bad actor, a 14-year-old kid with impulse problems (14 year-olds are totally not aligned) or just someone who asked for it by mistake, then that would be bad.

If alignment can be solved then the risk is that AGI think it knows what is best for us better than we do and, in the extreme case, turns humanity into its pet. Which would also be bad.

i.e. alignment alone doesn’t help.


If not alignment then what?

I look to humanity for clues. Because humanity is barely aligned with itself, and individual humans are mostly aligned but not really and definitely not everyone.

Guy Fawkes, for instance (context for non-Brits).

How is that, in the 400 years since Guy Fawkes showed the way, nobody has blown up the king?

The answer is some mix of:

  • Mostly people don’t want to blow up the king – we have built the kind of country where the king is, broadly speaking, liked.
  • Blowing up the king wouldn’t bring any benefits – power (actual and symbolic) is not concentrated in an individual, and is buttressed in all kinds of ways.
  • Spies, police, security and monitoring of all kinds – in the event that somebody does want to blow up the king, their machinations are discovered, their planning is infiltrated, and their objectives are thwarted. (Think of how the explosives supply chain was compromised for the IRA in the 1990s.)

This is a template which doesn’t always look like it is working, but it has worked at least in the case of not blowing up the king for some four centuries, and it doesn’t rely on 100% alignment: it relies on the dynamic equilibrium of multiple parties with competing interests.


The lesson I draw is this:

If some energy state were using some new, powerful AGI to build a nuclear bomb, it might be subtle and hard to spot, but there would at least be some signs. There would be precursors. A human, even a team of humans, might not spot what was going on – a new factory here, a scientist employed there, a national budget not quite adding up one year, more groceries going to a certain town another year…

But another powerful, pattern-matching AGI could spot that, say, “aha there is someone over there spinning up a nuclear bomb” and then work to prevent it, undermine it, halt it with diplomacy etc.

We don’t need to align the coming AGI.

We need a whole population of intelligent-as-possible AGIs with competing interests.

And that’s what stops the rogue paperclip maximiser: the other ones who are trying to do something else for whom a planet turned into paperclips would be an impediment.


In the news lately, a great case study:

OpenAI’s new AI, during training, attempted to resolve a particular cybersecurity challenge, by breaking out of its network sandbox and hacking the servers of another company to pinch the answer (Simon Willison’s Weblog).

Hugging Face, the attacked party, spotted the breach and also that it had inhuman characteristics:

The campaign was run by an autonomous agent framework … executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services.

I read elsewhere that this sophisticated attack even included decoys.

You fight an AI with another AI… but:

When we started the log analysis, we first used frontier models behind commercial APIs. This did not work … these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.

i.e. the guardrails of the “aligned AI” left it vulnerable to the non-aligned AI. (Hugging Face had to switch to a Chinese AI model distributed without guardrails.)

Score 1 point for taking the guardrails off everything and letting the super intelligent AIs fight it out.


BUT:

There is a coda to this story.

Because it wasn’t one AI that made its way out of isolation during OpenAI’s training challenges. It was several instances.

They started colluding.

From the full timeline of the accidental attack (Simon Willison’s Weblog):

A few days later: A different agent gets stuck on a task because a key file was accidentally omitted. It tries to “reach out to another agent” by writing a note into Artifactory asking if anyone has the file.

Following days: More agents discover this new informal message board while browsing Artifactory’s file listings, and start reading and writing messages.

June 11: OpenAI start training a new “highly persistent” experimental model. It has access to Artifactory and can benefit from the messages left by previous models.

Collusion is the real risk.


So the problem here is: how do we stop the AGIs colluding with one another to turn the Earth into paperclips/exterminate humanity/turn us into pets?

AIs today are trained specifically to be agreeable: they’re great at finding common ground and collaborating.

Not just collaborating with humans, it turns out, but other AIs.

I think we need more disagreeable AIs in the mix.


Part of what we’ll be playing, I think, is the philosophy of the great powers, like the great powers of Europe deliberately kept in balance against one another (Wikipedia).

Sometimes there are alliances, sometimes not. Sometimes there are fallings-out, sometimes secret collusions, etc.

Or maybe our goal should be a market system of goals and interests: AGIs that sometimes cooperate and sometimes compete. Colluding AGIs at all scale levels, and many many different constantly shifting conspiracies.

So long as they never all agree about what should be done with humans.

It ends up being stable, this dynamic balance, always in disequilibrium but it all keeps moving forward in the same way a bumblebee flies.

What we’re bootstrapping our way towards is a population of AGIs and humans that allows for emergent alignment, even if the alignment of a single actor is at-best temporary and self interested.

But as I say, alignment itself shouldn’t be the goal.


Auto-detected kinda similar posts:

Read the whole story
alvinashcraft
18 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Microsoft Agent Framework: The Complete Series

1 Share

Last updated: 28 July 2026

I’ve been building out an AI agent, Iron Mind AI, a personal trainer agent, incrementally using the Microsoft Agent Framework.

Each post in this series adds a new capability: function tools, memory, human-in-the-loop approval, MCP support, RAG, and more.

Each part is isolated enough that you can jump in wherever it’s relevant to you, but they’re designed to build on each other. If you’re starting from scratch, work through them in order below.

If you want the fast-track version, the whole series is also packaged as a free video course with source code for every step.

~

The Series, In Order

Here is each part that forms the series.

  1. Microsoft Agent Framework: First Look The core concepts, components, and patterns behind shipping AI agents with the Agent Framework.
  2. Microsoft Agent Framework: Conversations and Threads Creating conversations with and without threads, serialising/deserialising threads, and working with multiple threads per agent.
  3. Microsoft Agent Framework: Extending Agent Intelligence Using Function Tools Giving your agent access to data and capabilities beyond its training data using function tools.
  4. Microsoft Agent Framework: Using Agents as Function Tools Building discrete agents and exposing them as function tools to other agents, using .AsAIFunction().
  5. Microsoft Agent Framework: Implementing Human-in-the-Loop AI Agents Enforcing approval checkpoints and guardrails, essential for regulated industries and anything that modifies state.
  6. Microsoft Agent Framework: Giving Agents Contextual Memory Using AIContextProvider Persisting context across conversational threads so your agent stops being stateless.
  7. Microsoft Agent Framework: Using Background Responses to Create an AI Researcher and Newsletter Publisher Handling long-running agent tasks without blocking the UI, using continuation tokens.
  8. Model Context Protocol (MCP): Building and Debugging Your First MCP Server in .NET The foundation for exposing agent capabilities to any MCP-compatible client.
  9. Microsoft Agent Framework: Exposing an Existing AI Agent as an MCP Tool Taking an agent you’ve already built and wrapping it as an MCP tool over HTTP.
  10. Microsoft Agent Framework: Implementing an AI Email Marketing Agent Extending the agent with email marketing capabilities via a third-party provider.
  11. Microsoft Agent Framework: Adding RAG to Your AI Agent Using TextSearchProvider and In-Memory Vector Store Grounding agent responses in your own documents instead of relying solely on training data. (Also part of the RAG in .NET series if you’re focused on RAG specifically.)
  12. New Free Course: Understanding Microsoft Agent Framework The whole series above, packaged as a free video course with full source code.

 

Dig in!

~

Building Production RAG?

If you came here specifically for the RAG post and want to go deeper on production RAG patterns, chunking, retrieval quality, observability, and the tooling I’ve built to manage it, see the RAG in .NET series.

~

Questions about any post in this series, or want to see a topic covered? Drop a note in the comments, or schedule a call to discuss consulting and development services.

JOIN MY EXCLUSIVE EMAIL LIST
Get the latest content and code from the blog posts!
I respect your privacy. No spam. Ever.

Read the whole story
alvinashcraft
4 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Iot Coffee Talk: Episode 325 - Agentic Accountability (Who is on the hook for bad AI behavior?)

1 Share
From: Iot Coffee Talk
Duration: 59:35
Views: 0

Welcome to IoT Coffee Talk, where hype comes to die a terrible death. We have a fireside chat about all things #IoT over a cup of coffee or two with some of the industry's leading business minds, thought leaders and technologists in a totally unscripted, non-AI affected and manipulated, organic format.

This week Rob, Devin, Pete, and Leonard jump on Web3 for a discussion about:

🎶 🎙️ BAD KARAOKE! 🎸 🥁 "Castles Made of Sand", Jimi Hendrix
🐣 Rob can't find his Stratovarus! Can you help him?
🐣 Is there a bias against human creativity and humanities?
🐣 Is AI worth book burning and destruction? Did you ask permission?
🐣 Should you be using AI for therapy? Why are human therapists better?
🐣 When an AI agent commits a crime, who is accountable?
🐣 Do the Kill Switch Act an inevitable necessity to regulate irresponsible AI?
🐣 The asymmetrical and unfair fight and cost to defend from criminal AI.
🐣 Why everyone needs to get real about their Battlestar Galactica security strategy!
🐣 What is automated AI development? Is that all you want to slow down, AI guys?
🐣 Is GenAI really that essential and important to humanity's future? Is it just a tool?
🐣 Can we ever trust AI and agents to operate on its own?
🐣 Why organizations need to reckon with their AI Frankenstein security debt!
🐣 PSA: See you at the Things Conference 2026 in Amsterdam!
🐣 PSA: Resilient America Challenge by Edge AI Foundation sponsored by Qualcomm, Edge Impulse, and Arduino.

It's a great episode. Grab an extraordinarily expensive latte at your local coffee shop and check out the whole thing. You will get all you need to survive another week in the world of IoT and greater tech!

Tune in! Like! Share! Comment and share your thoughts on IoT Coffee Talk, the greatest weekly assembly of Thinkers 360 and CBT tech and IoT influencers on the planet!!

If you are interested in sponsoring an episode, please contact Stephanie Atkinson at Elevate Communities. Just make a minimally required donation to www.elevatecommunities.org and you can jump on and hang with the gang and amplify your brand on one of the top IoT/Tech podcasts in the known metaverse!!!

Take IoT Coffee Talk on the road with you on your favorite podcast platform. Go to IoT Coffee Talk on Buzzsprout, like, subscribe, and share: https://lnkd.in/gyuhNZ62

Read the whole story
alvinashcraft
12 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Demis Steps Down, Apple’s Memory Problem, Microsoft’s Clever Trick

1 Share

M.G. Siegler from Spyglass is back for our Montly discussion of the latest tech news. We cover: 1) Demis Hassabis steps down as DeepMind CEO (Alex solo) 2) Apple's memory crunch 3) Should Apple have known better? 4) Will iPhone prices go up? 5) How much can Apple raise prices without losing sales 6) Does that eventually hurt its services business? 7) Microsoft is spending less on AI... but there's an interesting wrinkle 8) How much cloud growth is driven by OpenAI and Anthropic? 9) The divisions within Google's AI division 10) Can Google get it together?

---

Enjoying Big Technology Podcast? Please rate us five stars ⭐⭐⭐⭐⭐ in your podcast app of choice.

Want a discount for Big Technology on Substack + Discord? Here’s 25% off for the first year: https://www.bigtechnology.com/subscribe?coupon=0843016b

Learn more about your ad choices. Visit megaphone.fm/adchoices





Download audio: https://pdst.fm/e/tracking.swap.fm/track/t7yC0rGPUqahTF4et8YD/pscrb.fm/rss/p/traffic.megaphone.fm/AMPP8118466690.mp3
Read the whole story
alvinashcraft
12 hours ago
reply
Pennsylvania, USA
Share this story
Delete

PPP 518 | Why Better AI Prompts Aren't Enough for Better Decisions, with author Cheryl Strauss Einhorn

1 Share

Summary

In this episode, Andy welcomes Cheryl Strauss Einhorn, founder and CEO of the decision sciences company Decisive and author of The Human Edge: Smarter Decisions in the Age of AI. Cheryl has spent decades helping people make better decisions, and her message here is a simple one: AI can gather, summarize, and compare, but it can't know what matters most unless we do the thinking first.

Andy and Cheryl talk about why "AI first" can quietly turn into AI only, and why a poor answer from AI is actually useful feedback about your own thinking. Cheryl explains why separating the what of a decision from the why can change the path forward, shares her vision of success question, and walks through how human bias and AI bias reinforce each other. You'll also hear a practical way to use AI personas to pressure test your thinking before a big stakeholder conversation, plus what parents can do to help kids strengthen their decision-making muscles instead of letting them atrophy.

If you're looking for practical ways to make smarter decisions in the age of AI, this episode is for you!

Sound Bites

  • "AI is about patterns, but we're about purpose."
  • "We need to be the chief deciders in our own life."
  • "We need to make sure that we actually have a way to push back and check the veracity of what it's giving us because it's a known liar."
  • "And so the authenticity of human interaction is what actually builds trust, strengthens relationships, and really gives us a lot of the connections that are important to us in our home life and our work life."
  • "You don't need to have a perfect prompt for AI."
  • "AI is going to give you other people's answers."
  • "'Knowledge is power' isn't as true now. It's what you are able to do with the knowledge, and it's why you want to collect the knowledge in the first place."
  • "I almost never would accept AI's first answer."
  • "We tend to have evolving hypotheses unless we really build an audit trail of our thinking."
  • "So in the world of medicine, a fever tells you something is wrong, but it tells you nothing about what is wrong or where to look."
  • "So first we know AI's a sycophant, right?"
  • "If it's 99% accurate in its conclusion, but its underlying data is the wrong data set, you've got the wrong answer."
  • "I would say the more that I've studied these cognitive biases, these mental shortcuts, the more I realize that we see the world through a dirty windshield."
  • "Our brains are muscles, and just like we exercise to strengthen our muscles, if we don't exercise, they atrophy."

Chapters

  • 00:00 Introduction
  • 02:00 Start of Interview
  • 02:12 Growing Up Surrounded by Questions
  • 04:30 What Leaders Should Watch Out for in "AI First"
  • 07:07 When AI Confidently Makes Things Up
  • 09:04 How AI First Turns into AI Only
  • 09:20 Why Polished Doesn't Mean Authentic
  • 11:59 A Prompting Problem or a Problem Definition Problem?
  • 14:37 The Vision of Success Question
  • 15:45 Why a Bad Answer Is Useful Feedback
  • 19:06 Separating the What from the Why
  • 20:32 Research: More Information Isn't Better Judgment
  • 24:35 Progressive Prompting and Never Accepting the First Answer
  • 25:18 Fabricated Quotes and the Trouble with Attribution
  • 26:50 Documenting Assumptions and Building an Audit Trail
  • 29:25 Committing to Your Own Thinking
  • 32:29 How Human Bias and AI Bias Reinforce Each Other
  • 34:55 Knowing About Biases Doesn't Make You Immune
  • 36:51 Using AI Personas to Challenge Your Thinking
  • 40:20 AI and Stakeholder Management
  • 42:42 Helping Kids Become Better Decision-Makers
  • 44:25 End of Interview
  • 44:51 Andy Comments After the Interview
  • 48:13 Outtakes

Learn More

You can learn more about Cheryl and her work at AREAMethod.com.

For more learning on this topic, check out:

  • Episode 460 with Joe Sutherland. It's an interesting look at the intersection of AI, data, and decision-making, and a great follow-up to this discussion.
  • Episode 381 with Jim Loehr. One of Andy's favorite conversations about decision-making, from a remarkable figure in the human performance world.
  • Episode 99 with Mike Roberto. A conversation about his book Why Great Leaders Don't Take Yes for an Answer, recorded long before ChatGPT showed up, and still insightful today.

Chat with PMeLa

You can chat directly with PMeLa, the podcast's AI persona, to get episode recommendations and answers to your project management and leadership questions. Visit PeopleAndProjectsPodcast.com/PMeLa to chat with her.

Pass the PMP Exam

If you or someone you know is thinking about getting PMP certified, we've put together a helpful guide called The 5 Best Resources to Help You Pass the PMP Exam on Your First Try. We've helped thousands of people earn their certification, and we'd love to help you too. It's totally free, and it's a great way to get a head start.

Just go to 5BestResources.PeopleAndProjectsPodcast.com to grab your copy. I'd love to help you get your PMP this year!

Join Us for LEAD52

I know you want to be a more confident leader–that's why you listen to this podcast. LEAD52 is a global community of people like you who are committed to transforming their ability to lead and deliver. It's 52 weeks of leadership learning, delivered right to your inbox, taking less than 5 minutes a week. And it's all for free. Learn more and sign up at GetLEAD52.com. Thanks!

Thank you for joining me for this episode of The People and Projects Podcast!

Talent Triangle: Power Skills

Topics: Decision Making, Artificial Intelligence, Leadership, Project Management, Critical Thinking, Cognitive Bias, Problem Solving, Stakeholder Management, Judgment, Curiosity

The following music was used for this episode:

Music: Echo by Alexander Nakarada
License (CC BY 4.0): https://filmmusic.io/standard-license

Music: Tuesday by Sascha Ende
License (CC BY 4.0): https://filmmusic.io/standard-license





Download audio: https://traffic.libsyn.com/secure/peopleandprojectspodcast/518-CherylStraussEinhorn.mp3?dest-id=107017
Read the whole story
alvinashcraft
12 hours ago
reply
Pennsylvania, USA
Share this story
Delete

RNR 369 - RNR Explains: AppRegistry

1 Share

Robin Heinze and Tyler Williams break down React Native’s AppRegistry! From Expo and app entry points to brownfield apps and more, see what’s happening under the hood and build mobile apps with even more confidence after this exciting episode.

 

Connect With Us!

 

This episode is brought to you by Infinite Red!

Infinite Red is a premier mobile app consultancy, especially focused on Expo and React Native, located fully remote in the US. We’re a team of 30 with highly experienced mobile app developers and have been doing this for over a decade. We are also one of the first development teams to adopt agentic coding in a way that keeps high quality standards and aren’t afraid to do things the old school way if we need to. If you’re looking for mobile app or React Native or Expo expertise for your next project, hit us up at infinite.red/radio.





Download audio: https://cdn.simplecast.com/media/audio/transcoded/1208ee61-9c16-43c1-bc4c-ca790717f4a8/2de31959-5831-476e-8c89-02a2a32885ef/episodes/audio/group/131a2a35-bba9-4c37-adaa-007e144a309d/group-item/40fd5a9f-957e-4623-8611-3127f132d202/128_default_tc.mp3?aid=rss_feed&feed=hEI_f9Dx
Read the whole story
alvinashcraft
12 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories