Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161002 stories
·
33 followers

Why the current tech backlash feels different

1 Share
A photo illustration of an email logo.

This interview has been lightly edited for length and clarity. 

Nick Statt: Hello and welcome to Decoder, Nilay’s show about big ideas and other problems. This is Nick Statt, senior producer. And I’m joined by our brand-new supervising producer, Greg Ott.

Greg Ott: Good day, everyone. And Hi, Nilay. Nilay is here too. He is the person who hosts the show, and his name is also in the show. So it makes sense.

Nilay Patel: It’s true. We don’t consistently say my name in the show enough. We should do it all the time. I should do it. Welcome to Decoder with Nilay Patel.

GO: Like Nick was saying, this is Nilay’s show and we are doing a mailbag episode. This is where we go through all the feedback. Because you listen to the end of every episode, we know you do. We do mention that we go through every single email that we get. Nick, you can verify that. We get a ton of email, and we do read through all of it. Is that true?

NS: We now get YouTube comments. We get many, many Verge comments. We get Spotify comments and posts on Bluesky, Threads, and elsewhere. So yeah, we get a lot of feedback.

I do feel like promising we’ll read all the YouTube comments is an emotional commitment that we’re making, but we do read everything.

GO: We do.

NS: I do read the Spotify comments. I read every one because those are specifically pretty pointed.

GO: Yeah. They’re kind of like uncut gems. Because they’re trying to build this little community, somebody might read this because it only gets like six comments as opposed to 600 for a YouTube video.

NS: We last did a mailbag episode in April, and we said we were going to do more of these. So we’re back again. This is September. We’ve had a few months. We’ve had a lot of really big interviews. So yeah, let’s just get into it.

GO: We’re going to kick things off with the “software brain” video you put out. This was one of the biggest things that people have commented on. It is far and away the most talked-about episode in the past six months. Hundreds of comments on The Verge website, hundreds of comments on YouTube, single to double-digit comments on Spotify, and lots and lots of emails. 

We want to start there because a lot of the feedback was positive and people actually rarely email us a ton of negative feedback. We get a lot of, “Attaboy.” Some guy I knew in high school just emailed me — Hi Dean — saying how much he liked our Kathy Hochul episode. So people like the show, but for this one, we got some negative stuff on The Verge that we’re going to get to in a second.

One thing we wanted to talk about though is something a reader brought up about the idea of the Nilay rants. And of you effectively pulling out your inner Dennis Miller and really giving us your take on something. For this one, one listener wrote, “Just listened to ‘People do not yearn for automation,’ and all I can say is the people do yearn for more Nilay rants.” 

So we want to know, what is the origin of this type of rant? And then how much of the show’s DNA do you think is built around this idea of you giving us the rant, and the idea of you giving us your hot take on something really, really well thought out, as opposed to the interviews we usually do?

I feel like I should make another riff on the “People do not yearn for automation,” but as Dennis Miller — just full rant.

GO: We’re going to have to pull out some really good old references.

The power of AI means Dennis Miller can deliver software right now.

GO: “Let me tell you something, babe. Sam Altman, the last time he was sitting in a Whataburger…”

Right. We could just do it. It’d be great. I’m sure we wouldn’t get sued for that.

They’re just essays. I know they’re rants and they’re delivered on YouTube and they’re very pointed in particular ways. But that’s just what I would’ve written as a column 10 years ago before we had podcasts this way. It has been very obvious to everyone that you can’t just write a column and expect to have an impact. You have to stare at the camera and do it. We’re very good at publishing that as a column and then having a transcript and polishing it and delivering it all the other ways. 

That’s the origin of it: For all the podcasting I do and as little writing as I do anymore, in my heart, I’m still a writer. Taking all of the reporting and all of the ideas that we hear and all the conversations we have and synthesizing them into ideas that you can grapple with is the point. It is really the point of running a magazine.

That’s where that comes from. Now, can I do it as often as I want to do it? I obviously can’t. The reason I did “software brain” at that moment is that the excitement around AI is so high. The hype is so high and it’s all connected to how good AI systems are at writing software and manipulating databases. But the reason that I thought, “I got to make the time to do ‘software brain’ right now,” is that I could see the hype cycle around AI failing to contend with the lack of verifiability in every domain except software. That’s getting a little blurry now.

Lots of companies are just databases. Software engineering is verifiable, you can run it through the compiler and see if it works. AI has obviously gotten really good at making software. If you believe the world is just software, that means AI has gotten good, it’s done. You can see the labs are talking about how they’ve achieved AGI or they will very soon and it is very much connected to the ability of AI to write software. I’m just not discounting that that’s important. I’m just saying once you get outside of software, verifiability gets way harder.

You can see it in the episode we did with Verge reporter Robert Hart about AI and math. There’s an amount of verifiability in math and you can burn a lot of tokens, we don’t know how many, and solve some unsolved math problems. There’s probably verifiability in various other science domains, so it’s really interesting. Is there verifiability in generating novel drugs? No, you got to inject a bunch of people with novel drugs. What are we going to do?

GO: That’s why they got so many rats in cages somewhere, it has to be done.

Right. There’s just some other thing that’s going to happen in the next 10 domains that is going to be really, really challenging. That’s where I wanted to put the marker down and say, “I have been swimming in this water for a long time with these people and the power of software is not to be denied, it has remade our world. But if that’s the only framework or you think that is a universal framework, you’re going to run into these walls.” 

There are lots more of those we can do, but I want them to hit, I only want to write hit singles. Importantly, they’re all born of the reporting. That’s always been the dynamic on The Verge. Certainly the dynamic on Decoder is the data can get to your head.

The episodes where I just talk sometimes outperform the episodes with CEOs. That’s not good. That’s a bad incentive loop for me because I want everything to come out of the fact that I’ve talked to all of the people and you can go check my work. You can go see what those conversations are like and you can see that they are genuine and honest in all the things they should be.

That’s where I feel like I not only have the depth and the foundation to write something like “software brain,” but I can send it to those people and say, “This is how I feel about it.” They come on the show and they talk to me about Google Zero or “software brain” and all that stuff, and that conversation gets much bigger than me just ranting.

NS: My reaction to that, Nilay, is that you feel like your thesis for “software brain” has borne out over the last few months. We did get feedback from people who disagreed with some of the premises.

It wasn’t that they’re saying people aren’t unhappy with AI, but they think that we’re wrong about why they’re unhappy with AI. One reader, Josh Vanderberg, wrote in to say that he thinks it’s not that people hate AI as a technology, but rather that they don’t like their current economic situation and they’re projecting that onto AI. He’s pretty sure that we’ll eventually embrace natural language as the superior way of doing things on a computer.

He added, “All modern computer interfaces suck hard. We’re forced to map our fluid non-deterministic creative selves into hierarchical menu systems. No matter how intuitive and discoverable, we constantly fail to find the tools that we need.” He uses Canva’s CEO as an example of somebody that deeply understands the problem and they realize that AI can provide the user interface we’ve all been yearning for. He goes on to say, “People definitely yearn to just tell Photoshop what they want done in words and forget everything they ever learned about that horrible user interface.” 

We had a few other similar comments like, “People do not yearn for automation, they yearn for cheap stuff. And when AI gives them cheap stuff, they’ll be happy.” A third person wrote, “’Software brain’ misses three things regular people actually care about: control, agency, and meaning. We like doing things. The doing is where the meaning comes from.” 

I’m curious just from those particular comments and also the other feedback you got from “software brain,” and how you respond to those criticisms?

Let’s take them in turn. This is a very thoughtful comment, Josh. I don’t disagree that all modern computer interfaces suck hard. I’ve used Liquid Glass. I agree with you. I refuse to upgrade this computer to macOS Tahoe. I’m just not doing it. 

There’s some confusing thing happening in user interfaces, probably related to how much all of our devices have turned into shopping malls or SaaS products or whatever it is. There’s something broken there and I agree that that’s bad. I disagree that natural language is a perfect substitute for everything or can become a universal interface.

I say this as somebody who is writing a Decoder book based on all the interviews on the show. I’ve used Wispr Flow to get as much out of my head onto the page as I can. That is incredibly powerful and you can see why people like it. I also have to edit every single word it generates because it insists on formatting my thoughts in ways that I find completely bananas. All these systems say, “You think in bullets.” I say, “I am not thinking in bullets right now. I’m just pausing between sentences.”

There’s something there where you can take the natural language input and what we as people are reasonably good at doing is taking everything else with it, the context of our body language, the tone of our voice, and creating other meaning out of it. Computers are really bad at that. Maybe they’re going to get better at it, but they’re currently very, very bad at it. I’m using my Wispr Flow example. It’s just one that you can see how bad it is at that.

It’s not just you’re being creative to a computer or talking to a computer that’s going to do what you say and you don’t have to push all the buttons, it’s lossy. It’s lossy in a very important way that using a mouse to click on a button is not lossy at all. It is very, very directed and it will do exactly the thing that you want it to do. Ordering from Grubhub with your finger on a phone screen is 100 percent accurate — unless you’re drunk, which I’ve been. 

But talking to an agent to have it to go do something is lossy. It is error-prone in very specific ways. These things are going to exist together. That, to me, is the balance that we’re going to have to strike, that all of the hype around AI assumes that what I want is another person instead of a specific thing to happen. You can talk to lots of designers or professionals who use these tools and they will all complain about these tools. 

We use this app called Riverside to produce the show several times a week. Every single person on this team can issue a PhD thesis in all of the problems with Riverside because it’s the tool we use and it sucks in very specific ways. It is also still the tool we use every single week to make this show. If anybody uses Photoshop or Pro Tools… the books that can be written about people’s feelings about Pro Tools could fill a library, but they’re the tools because they get the job done in specific ways.

Making that lossy or making that non-deterministic is actually going to make people less effective. It will democratize access for a lot of people, I think, and that is a huge opportunity. I don’t think that opportunity outweighs the cost of making things more chaotic in specific ways.

The cheap stuff piece: yeah, I think people like cheap stuff. There’s a reason TVs are permanent surveillance devices and everyone has them in their home because they cost $500 and they’re just watching you all the time. But I actually think the industry is coming around to my point of view that the products aren’t good enough. 

I know this because our friend Alex Heath just interviewed Sam Altman and he asked him why people don’t like AI. Altman’s response was that he thinks that the way to win hearts and minds is to deliver value. We actually have a clip, you can run it.

Sam Altman: Generally speaking, I think the right way to get people to like something is to deliver them value… Probably if a lot of people use [agents] — which they will over time — and understand that it’s not actually fusing and destroying huge amounts of water or whatever, then they’ll be more excited.

I watched that, and I think Alex did a good job talking to Altman. My view of that is it’s a concession that Sam Altman knows ChatGPT is not delivering enough value to consumers. These free versions are not communicating enough what they could do or they’re not doing enough or they’re hallucinating too much and that’s why consumers dislike them.

I continue to believe that talking about how everything can be turned into a database that can then be automated by agents is not going to win anyone any hearts and minds. Maybe that’s the technology that creates products people love, but the technology itself is fundamentally uninspiring. It’s the products, it’s the experiences you have that change minds, and “software brain” is inherently alienating, in my opinion.

The last one, “’Software brain’ misses three things people care about: control, agency, and meaning.” Yes, I agree with that. I would connect that actually to buttons in cars. We’ve gone through this horrible period where all of the controls in cars became virtualized onto a screen. There’s been furious pushback against it and all the buttons are back. I think people want deterministic actions, even from computer-like things. 

This is maybe the single best market-driven evidence that you can’t just virtualize everything. You actually have to give real experiences to people that do the things that are printed on the label of the button exactly when you push the button. That is control and agency. I don’t know about “meaning” in changing an HVAC in a car, but it’s definitely control and agency. Meaning I think comes from having to struggle through doing the work yourself, not just letting AI do it for you.

So I agree with all these comments in different ways. I just think the industry has gotten completely lost in the sauce on what you can do if you turn everyone into a database, which is what I call “software brain,” and they’ve missed the human connection. A really good example — I’ll give you a little preview of an upcoming episode — we had the CEO of Cloudflare on, and he just laid off a bunch of people at his company because he says AI can do a bunch of those jobs.

His whole thing is he can identify rising talent in the company by having the AI look at all the things they’re doing. Then he thinks that’s good because he can make a phone call and say, “Hey, you’re a young superstar.” Maybe that is a good thing. In order to accomplish that goal, every single thing that happens to that company needs to happen in a database for an AI to review and then communicate to him. That’s the preview. We had a whole conversation about that because I think that is one of the more fascinating ideas in all of business right now.

GO: Boy, I hope that doesn’t roll out to The Verge. I’ll have to get one of those mouse movers or something. You’re going to have to start figuring out what I’m doing all day.

We are maybe insistently human. If we could automate Riverside, I think that would be good.

GO: Yeah, good luck. They did roll out the option to build an AI twin of you. The features we’ve been asking for — not better microphone connections, not better uploading. No, we need to duplicate the people we’re speaking to.

Yeah. Can we turn off the auto level control? No. But AI twins are here.

GO: There’s actually a request for an essay-style thing, a video essay that is a riff on the “software brain,” titled “People Do Not Yearn for Surveillance,” back to the whole watching-what-you’re-doing-at-work type of thing. 

Is that something that’s in your head right now, especially with all the Flock cameras everywhere, all the controversy around that? Maybe with a “software brain” format or framing or just a bigger picture thing? Is this in your mind right now?

It is a little surprising to me. I’ll candidly admit, I wasn’t expecting it to be as bipartisan as it is. Some of that is just that it’s an election year and people are angry about cameras, so all the politicians are trying to channel that anger, maybe the same way they’re trying to channel anger about data centers.

But we’ve seen there’s action in red states to ban surveillance as well, and there’s action in more Republican communities to take down the cameras, which I think is fascinating. It’s not just limited to brands like Flock, it’s the idea of these cameras watching. 

The piece of the puzzle, the reason that it caught me by surprise, is that for better or worse, my formative political experiences were 9/11 and the war in Iraq. That’s when I was in college. And what did we do? We passed the Patriot Act. We militarized the police. We started doing warrantless wiretaps. We ran our way all the way up to the Snowden revelations and then we did nothing about them. Literally none of that stuff caused this amount of panic. That’s just my framework.

This is a shock to that system. It’s surprising to me that people care now. I think it’s because the cameras are visible. The understanding that there’s an AI chatbot that you can talk to that can just go generate an answer and the cops now have it for cameras that are tracking you everywhere. And you can literally see the cameras and you can connect it to an experience you’ve had.

There’s Flock, who we are trying to get on the show, and their CEO originally tried to do a bunch of MAGA-coded messaging around this stuff. He would call his critics terrorists and that didn’t work. All that is new if you are born of the things that I was born of. All of these are criticisms of the Patriot Act in 2001, and we just did it anyway. I’m curious for the feedback from the audience on what they think shifted that made a bunch of stuff that we were happy to do in the past feel like one of the more potent political issues of our time?

GO: I tie this to the Meta glasses right now with the way those cameras are being received. Compare that to how Google Glass rolled out. If you refer back to the Patriot Act days and your formative years of tech, that’s back when cellphone cameras were seen as a joke. “Wait, I’m going to take a photo on my phone?” Right now the only thing I would take a photo with is my phone because it’s the best camera I will ever have with me. 

I find that to be a really interesting evolution, especially tied to the AI of it all. When it’s like, “No, no, these cameras now can connect to this network and have all this extra information.” It’s not just some file living on your device that you can put somewhere with your discretion. No, it’s out of your hands now. I wonder if part of that’s the anxiety around things like the Flock cameras, the glasses cameras, and all that.

I definitely think there’s some amount of anxiety around “I’m being recorded without my consent.” There’s a lot of anxiety around the intent of the people particularly wearing Meta Glasses. There’s a reason they’re called pervert glasses. Some of the biggest influencers in the world have posted videos complaining that they get stalked at the airport and people wearing Meta Glasses film them coming off of planes. This is very bad. 

It was Olivia Dunne who posted a video about that. She specifically called out Meta Glasses, which is not a great rep for those products to have. There’s something about that. There’s something about how intrusive the phones are and the cameras are and the idea that everything is recorded and meant to be shared, but we’ve been living in that for a long time. Our thesis there is that there’s a button on a phone and you pushed it and you kicked off this line of huge consequences because the camera exists in your pocket, like you’re saying, Greg.

I’m just saying there’s something that happened recently where the very notion of the police having a network of cameras that are all connected to each other and they can do an AI-powered search has just flipped over. Now people say, “We’re going to go and cover them up with spray foam or we’re going to tear them down.” Or 404 Media, which is excellent, has done so many videos finding clips of people at town hall meetings yelling about Flock cameras and being shut down by their own elected officials.

There’s something there that’s new. We’re obviously going to cover it. I’m not saying it’s so surprising we can’t engage with it. I’m saying from my worldview, the dirt that I was raised in, the thing is the same as it ever was, and that’s why it was surprising to me. We could have had this conversation about the Patriot Act. We could have had this conversation about the existence of the TSA, or any number of things, and we didn’t. We could have had this conversation about the Snowden revelations. Something flipped and to me it’s cause for investigation and also reflection.

NS: That’s a good time to switch gears to talk about our most popular episode of the year. That was, of course, our sit-down with Google CEO, Sundar Pichai. This was our fifth year in a row talking to Pichai at Google I/O, and it always generates a lot of feedback. 

This year we had commenters write in to say, “There seems to be an underlying answer of, well, we have to do all of this even if it’s bad or irresponsible.” And that person also said, “It’s a bit odd talking to him without getting some stronger concessions from Sundar.” Another said, “This interview is great, but I think Sundar got away with not having to answer the root issue of Google Zero, which is in this example, how Conde Nast were seeing a reduced click-through rate because of AI summaries.”

The question here for you, Nilay, is: Do you feel like Google is getting away with not being held responsible for what it’s doing to the web with AI? And do you think Sundar Pichai has a sense of Google’s culpability for the swath of destruction that’s being left in the wake of the AI era, and do you think he’s ever going to address or acknowledge this? Or is he still going to disagree with the premise five years from now?

That’s a lot of questions. I’m going to just issue the caveat that comes up in every single mailbag episode. Sometimes it feels to me like the audience is mad that I don’t arrest the CEOs on Decoder.

GO: [Laughs] We’ll arrest them over Riverside. Yeah, that’s a future we don’t have.

Maybe one day. Well, I was in person with Pichai. We could have cuffed him.

GO: That’s true. That’s right.

Maybe one day I’ll be granted this power. I’ll be deputized by the internet police to arrest the CEOs. But all I can do is ask the questions, and it has been five years of talking to Pichai. I hope he comes back next year. It’s remarkable that he’s the one who comes and takes the heat. We ask them all, and I can’t make them show up. They have to decide they want to do it, and he decides he wants to do it. I appreciate that he shows up and takes the heat and that he’s willing to do it.

Let’s use the Conde Nast example. I asked that question before a bunch of publishers started saying to other media outlets that they were going to stop letting Google index their sites. That’s the last turn. When they are so clear that it’s not worth it for their content to be stolen by AI because they’re getting no search traffic, that they just turn off the Google indexer, the Google crawler entirely. 

My conversation with Pichai would have been different if I had had that in my back pocket. I’m looking at these reports from various big publishers like USA Today and People Inc. And they’re all the way at, “We’re going to stop letting you index our sites because the traffic isn’t coming.” I don’t know what he would have said. I just think that answer might have been different or that the tenor of that moment might have been different.

What I have been taking from Sundar Pichai these past five years, and which I agree with, is the sense that these publishers got addicted to Google and they didn’t do anything about it. The shape of that conversation or his frustration, if you ever sense the flash of frustration that I sometimes sense is, “The web is really big. Google can measure the web. They have Chrome, they have all this data about the web. And the web is getting bigger and some parts of the web are really, really vibrant.” 

If you listen to Amy Lanzi, the CEO of Digitas, when she was on Decoder, she said, “We are getting more calls from our clients, more requests to make webpages and landing pages than ever before.” So some part of the web is growing and Google has that data. I think Pichai is saying, “These publishers are constantly complaining that I’m taking their business away and they did nothing else about it.” Every time I talk to a publisher, I ask them, “Why didn’t you do anything about it?”

I agree with that. There’s some sense that Google bears the responsibility for making the people who built a business on their platforms successful. Every platform eventually kills every business that’s built on it. YouTube will eventually kill YouTubers. Instagram will eventually kill Instagram. If you’re a Decoder listener, you know I think this. I’m always asking creators what they’re going to do about platforms and platforms when they’re going to kill creators. The web is just a platform for Google. Do I think that Google’s getting away with not being held accountable? Sure, a little bit. But the mechanism of accountability is not letting Google index your site. 

I’m very curious to see if this handful of publishers who have started making noise that they’re going to de-index their sites actually go through with it and actually impose the leverage on Google that I feel like the company has been daring them to impose. Again, I don’t want to preview the Cloudflare conversation too much. There are other mechanisms and there are other ways of doing it and that was part of that conversation as well. 

But there’s something here where this straightforward “you were sending us traffic for free and we were giving you content for free” exchange is gone. I don’t know that you can provide accountability such that Google doesn’t just have to still compete with Anthropic and OpenAI. They have to.

GO: Speaking of publishers and web traffic and all that, do you need to drop in a very nice little disclaimer about our new company and the PMX situation?

Oh, this is so complicated. So through a series of what only can be described as divorces and remarriages, Vox Media has split itself in two and The Verge is now part of something called PMX Global, which is a subsidiary of PMC, which is the Penske Media Corporation. PMC is suing Google, I believe, over AI Overviews and traffic. PMX and the thing formerly known as Vox Media are all in ad tech antitrust cases against Google.

I’m telling you this because it is a disclosure we should make. You can tell how much I know about them. They’re happening very far away from this newsroom, but it is true that our companies are involved in active litigation with Google about all this stuff.

GO: Thank you. Just want to make sure we get that out there because you flagged this the other day for when it’s relevant and it feels like with everything you just said makes that quite relevant.

Yeah, it’s in the strike zone, but it has nothing to do with us.

GO: Pivoting to something different, we had a listener talk about how our AI skepticism broadly on Decoder is Western-centric and it’s pointing out how a lot of AI adoption in China is actually different and faster and more welcomed in other ways. 

This is from Alex Wong: “Have you considered Asia and more specifically China where, by outward appearances, AI rollout adoption and image seem to be faster, deeper, broader, and more well received than in the West? Do you find this to be true? If so, what is the difference? And if so, why the difference? And if not, what are we not seeing?” 

What are your Asian sources seeing and saying? Should you have more Asian sources? Is it time for Decoder and The Verge to expand beyond Silicon Valley? This is a US-based show. Nilay is based out of New York. The production team is scattered, but we have New York offices where people come and shoot stuff. 

Do you think there’s room for more non-Western perspectives? Can you speak about the differences that are evident between the way China is approaching AI development versus in the United States? And can we really cover such a massive transcontinental gap like that? It’s a lot.

I do want to point out, as you said, we’re in New York. I actually think we need more Silicon Valley perspectives. We can be myopic. New York is all about money. To be successful in New York, you have to be like, “Literally where’s the money?”

San Francisco cares less so about the money, and more about the big ideas. Notably, the AI companies are not making money and running into some harsh realities here in New York from time to time. So there’s that.

I am not going to speculate too much on my understanding of Chinese culture. I will just offer two guesses and they’re just pure guesses. One, we can only know what the Chinese state media wants us to know. There are speech controls in that country. And so whether or not there’s massive discontent is not something that would ever come to me.

That’s just a gap and I’m cognizant of the gap. I agree that what we’re seeing there is a bit more excitement. We’re seeing a bit more investment. I would connect that — and this is my second guess — to there being literally a much stronger social safety net in China. So the idea that this will be a threat, like an existential threat to your job or your career or your livelihood or your health insurance or whatever it is, may be ameliorated by that stronger social cohesion. That is a guess, and I really hesitate to overthink that.

I would love to hear from people with closer perspectives to that, but those are the things that I’m always thinking about. I have said to leaders of big AI labs in this country, “People would like you more if you would just pay for health insurance. If you could just find a way to support Medicare for All with your excess profits and your trillion-dollar TAMs, maybe people would shut up.” And they’ve all said, “Interesting.”

I would just offer that connection because maybe the best thing AI could do for everyone is make it easier to start a business and to make it easier to have agency over your own life. The number one reason my friends with children don’t start businesses is because they need health insurance. That’s the connection. Again, that’s a very Western perspective, a very US-centric perspective specifically, but that’s why I went directly to the social safety net. Maybe you’re more excited about this thing that lets you do more things because the risk is limited in some other ways by the very nature of the country you live in. I’m open to feedback on that.

I’m responding to the question with my series of guesses and assumptions. I agree. We do now work for a much bigger company and hopefully that means I can put some reporters in China or in the region and we can get some firsthand reporting. But I’m curious for feedback from the audience because of those two things: One, it’s hard to tell what people actually feel given the speech controls that exist in that country; and two, is it just a different sense of possibility because the risks of existing are different?

NS: Do you think Sam Altman actually supports universal basic income? Is OpenAI going to fund UBI?

I don’t know the answer to that question anymore. They’ve all learned to stop issuing pronouncements about the future of society as loudly as they had been and spend way more time focused on enterprise adoption and software engineering and agentic marketing solutions and potentially “here’s how everyone will become rich,” not “here’s how everyone will lose their job.” So I don’t know the answer to that, but UBI is one of those things that it’s an idea you have because you think no one will have a job. I think they’ve all needed to stop talking about that.

NS: That actually ties very neatly into something we wanted to end this particular segment on, which is a pretty prescient email from a reader, Toms Bernards Callahan, specifically around tech backlash. He wrote in, “The algorithmic fragmentation of content consumption disrupts any organic large-scale organization. Yes, there are rallies, but no pointed policy pushes. Feeds are purposely placing users into niche consumer groups and the news has a hard time permeating across these groups.”

He says, “The irony here is that algorithmic fragmentation will make the backlash worse because when the dam breaks, the masses will be quite angry and it’ll be a large cohort of people.” He wrote us that email in April and I do want to say it feels like something has definitely happened. It feels like something definitely broke over the last three to four months since we got that email from him.

It feels like there is a large bipartisan cohort of angry people. We see this, of course, like Florida and Texas banning Flock cameras, the upcoming Florida governor’s race being about data centers. What do you make of this, Nilay? Specifically this moment of tech backlash that we’re experiencing and how The Verge might cover this from the perspective of algorithmic fragmentation? Or the perspective of this issue creating a bipartisan coalition that could kind of permeate through your filter bubble or your algorithmic feeds? 

The thing that really strikes me about the backlash right now is it’s a backlash to imposition. Not to speak for all Americans, but Americans don’t like being told what to do. “You have to accept these data centers, even though you’re telling your local governments you don’t want them” is an imposition that is just easy to be opposed to regardless of your political alignment. 

These cameras are going to watch everybody all the time and we have records that cops are watching kids in gymnastics centers. It is the most incredible imposition of all time. The nation was founded to rebel against this kind of imposition. It’s in the Bill of Rights. Of course that’s going to overcome algorithmic feeds. You can’t turn the knobs on TikTok enough to overcome that. No one can hear a fake AI influencer talking about how much they love Flock cameras and believe it’s true. It’s just not in our national character in a very specific way.

That’s like you telling everybody what’s going to happen to them and its power being imposed on them or in the case of data centers being made more expensive for them, it’s natural that that would overcome some algorithmic fragmentation. The other issue is it’s an election year. We see politicians scrambling to still play culture war games. There’s Benn Jordan, who’s a musician and a researcher who’s done really incredible work on Flock cameras. We interviewed him on The Verge before. He was just on the Tucker Carlson Show. Their views do not align. Even Benn said, “I can’t believe I’m here.” They were there to talk about Flock cameras and surveillance. That’s bananas, just purely bananas.

There’s something about that that I would connect to. At least with other big tech impositions, it felt market driven in some way. Facebook is going to be big and you’re going to be mad that Mark Zuckerberg is listening to you, but you have a desire to use Instagram. And boy, does it turn out that people’s revealed preferences are for Instagram to keep happening regardless of whether or not you think the ads are listening to you. 

There’s no market preference that’s going to take that away from you or there’s a network effect so strong, something else will happen and then we’ll try to regulate it. The market is still going to speak. People are still going to want to use this thing. That’s not true of Flock, there’s no consumer use for Flock. It’s diffuse. You have to prove the crime went down somewhere else later, but you can see it right now and it’s watching you right now and the cost is available.

AI has mostly been imposed on consumers. Every time I open Google Meet, it insistently tells me Gemini is here now. I don’t need this to be here in any way, shape, or form. That imposition where we’re claiming all the usage stats are so high, when in many cases it has just been put in front of us — there’s something about that that feels like it can transcend the algorithm.

GO: There’s something to what you were saying a few minutes ago about the Patriot Act of it all. There’s such a lack of legislation on a lot of these important issues that people feel like they’ve had no input on because they’re just rolled out. And the message is, “You should just take it.”

To your point, these people who you wouldn’t expect to go be a guest on Tucker Carlson’s program, it is this scrambling of an alignment because people don’t feel that agency. Nobody’s representing them on their behalf in a lot of ways.

Yeah, it’s election year. What are you going to do about the fact that Donald Trump is running around saying that data centers are money printing factories and you’re stupid for opposing them? 

GO: He called data centers the golden goose. “That’s the golden goose. Does everybody want to be poor?” That’s what he said the other day.

I certainly cannot predict elections and how they’ll go, but that just seems totally out of step with everybody. There’s something about that that I find utterly fascinating. To me, it just comes back to that Americans don’t like being told what to do. If you put yourself in that position, boy, do a lot of things sort of fall into place.

GO: Keeping that Gadsden flag in the air, I’ve got a transition here to your interview with Kathy Hochul, who was on the show last week as of recording time. It was received this way. One listener wrote, “Great interview job, very pointed questions. Little surprised that she didn’t seem to have thought through the ID verification implications these social media laws will lead to.” 

Some other Hochul feedback was pretty negative, predictably. Every single answer she gave could be summarized with, “It’s a campaign year. This interview was unbearable. Politicians don’t give answers to questions, they just give self-aggrandizing stump speeches.”

You have said you are honing your politician interviewing skills. What was your takeaway from this whole experience with her and how do you think it should inform how Decoder interviews politicians? Do you think we should have more of them coming into the midterms and the 2028 election? How do you see the whole politics of it all?

We definitely need to have more politicians on the show this year. I’m very sorry. But as we have been discussing, the tech policy issues in this election cycle are almost overwhelming. There might not be other issues. There’s the war in Iran, there are data centers, and there’s surveillance. Somewhere in there is healthcare. Those are the ones, at least in my opinion. 

As we talked about with Governor Hochul, a lot of them are connected. Surveillance and AI are deeply, deeply connected. Age verification and surveillance are deeply, deeply connected. It’s funny, some of the feedback we got on this episode was, “Facebook knows so much about you, why can’t it just tell you your age?”

With surveillance, it’s AI and it’s age verification. It’s all just a remix of the same ideas. How much of your computer should someone be monitoring to allow you to do some things? That’s the 3D printer debate in that conversation. I suspect we’ll end up talking to more politicians because the opportunity for realignment is so high. 

I honestly hope we get some conservative politicians to talk about data centers and surveillance because I would like to understand the realignment from that perspective. That said, I don’t think I’m great at interviewing politicians. I think I’m good at interviewing founders. I’m good at interviewing product CEOs. I’m bad at interviewing McKinsey CEOs. I’m getting better. And I’m worse at interviewing politicians. The amount of media training, it rises. That’s the thing that I’m describing here. We’ve always described the show as “Nilay versus media training” and I’m always trying to fight to a draw.

I got a draw with Governor Hochul. We got her to concede that she was out of her depth on some of these issues, particularly around age verification. I thought the moment where she said, “Tell Cody we’re going to get ahead of him next on 3D guns,” was very revealing because it’s a great spiky cowboy thing to say. I mean, that’s the moment you want from the interview with a politician. You want the sound bite. But it just struck me that that was a moment about conflict. It was not a moment about what people in our comments on the transcript were bringing up, which is whether or not most people can even print a plastic part that can withstand a bullet being fired.

I thought that interview went well from my perspective, because it just demonstrated that it’s easy to talk about what tech policy should be, but it’s very hard to talk about how you would actually implement the details of that policy. We did a good job there of even illustrating for the governor where the gaps were because she conceded it several times. I understand it’s frustrating to listen to and maybe one day I’ll just be a full cable news anchor and just repeat the same question 10 times in a row. But I don’t think anybody wants me to do that. Cable news anchors exist and they do a good job of that thing. I’m trying to do something else here.

NS: On the topic of interviewing styles and techniques, Nilay, we had an interesting interview last month with Bluesky CEO Toni Schneider. One aspect of that interview that really touched a nerve was when you asked him what he thought about Bluesky’s reputation as a liberal bubble. This is what he had to say. 

Toni Schneider: Yes, we definitely want that to change. It is already changing. It certainly wasn’t designed to attract one specific group of people. It works for anyone… We’re an open network for anyone and we don’t want to serve just one type of user base. I’ve seen it already in the last few months, really trend in the right direction where we just have more and more people show up and build their own corners of the network. It sort of balances out over time. 

This inspired a lot of conversation, not just in comments and emails, but also on Bluesky itself and the replies to The Verge‘s post about this. But one listener wrote in to say, “The idea that this one specific, very active group of people is an imbalance to be straightened out rather than the very reason the platform is what it is, is the mistake that Bluesky will come to regret. What a shame.”

Some of the feedback on Bluesky itself was predictably harsher. Somebody said, “Go ahead and attract other people, but the second we start seeing fascist stuff in our feeds that we haven’t subscribed to, you’re going to lose your base.” Somebody else said, “Toni doesn’t even understand Bluesky’s own user base.” Another person said, “Cool. I’ll just delete my Bluesky like I deleted my Twitter.”

Looking back at this exchange, it does feel like he artfully dodged it. It’s hard to tell what he really thinks about Bluesky’s reputation and the liberal politics of it all. In moments like that, Nilay, do you think it’s better to push people for a follow-up to be clearer on what they think? Or is that a case where the backlash became evident after the fact and it was better that you just kind of let him say his piece and you weren’t really trying to bait him into any one direction?

I wasn’t trying to bait him into any one direction and I understand the backlash from the Bluesky user base. I also understand what every social product CEO is struggling with, which is that you have to grow and you can run out of people. You can run out of people who want to be in that bubble. You can run out of people who only want to talk about one sports team. You can run out of people who want to be in whatever bubble X has become. We actually are seeing X run out of people who want to be in that bubble. The platform is shrinking. 

You can describe that as the capitalistic urge to grow, but the platform CEOs also understand that to remain vibrant, they have to constantly be attracting new people. Then you almost have to issue an answer like this, which is, “We’re a platform for everyone.” Toni’s riff on it was, “And we’ll give you the moderation controls to craft the experience you want to make the big platform feel small.”

I don’t know if they’re successful in communicating that to people. The big disconnect with Bluesky, in my opinion, this whole time, is that it was started by a bunch of ideological-technical people who wanted to build a standard called AT Protocol to enable interoperable social networking. Because of X, they accidentally got a user base and they’ve been in conflict with that user base ever since, which is fine. It’s obviously reflected in these comments, but Toni’s job is to make the thing a product. That’s why he’s the CEO. These answers just track every other social product answer we’ve ever gotten.

I don’t know if he’s okay with the bubble he has drifting away. I know that we can see in the stats that it’s getting smaller. I’m sure there’s going to be some revolt when they monetize it in ways that he’s described monetizing it. But if you’re a social product and you aren’t growing or trending towards growth, then in fact you are dying. There’s something about that that will not allow you to persist. That is a pattern that’s pretty repeatable and pretty understandable.

Maybe because we’ve spent so much time talking to social platform CEOs, I knew what he was getting at. I’ve internalized that answer in many different ways. I probably should have made him say it out loud. That’s a good note for me, but what I was hearing at that moment was, “You have to grow and that means you have to appeal to many more people.”

GO: Speaking of appealing to people, we want to appeal to our audience. They have some more thoughts on the show Decoder in and of itself, not just wanting more rants from you. We’ve got feedback from listeners who want to know about the kinds of relationships that you and we have with the companies that are covered on the show.

This is from Len Art who emailed to say, “Is there a possibility for you to make an episode about different ways the companies try to pressure The Verge into editorial changes? I don’t mean you to out all the company names or people making the requests, but just hearing concrete examples of how the requests come in and how they’re being dealt with could be a very compelling episode.” That’s what Len thinks.

What do you make of that? Do you worry there’s such a thing as too much transparency as we make a thing like this? There’s a lot of discretion and nuance that goes into negotiating these conversations and making sure that people want to talk to you. It’s a lot of back and forth. 

We are old enough and famously spiky enough that this stuff doesn’t happen very often. Yes, there’s the idea of too much transparency. My worry about too much transparency is that it would just be boring. We would say, “Nick and Kate spend hours scheduling.”

That’s actually the hardest part of making the show for us at this point. It’s CEOs, their comms people, and whatever news cycle they want to be on. It’s my schedule, which is bananas. Honestly, I’m telling you, and you can believe me or not, that might be the most complicated thing that we’re doing. We only make so many episodes of this thing a year and a lot of people want to be on the show. We’re lucky. That’s very lucky.

There is pressure. Big tech CEOs don’t want to be on the show. It’s not that we don’t ask them, they just are pretty honest sometimes that they think it would be too hostile of an environment. They want to be more in control of that message. You can see where they go instead. Do I want to play that game? I don’t. You can see that we just haven’t. We’re pretty direct, but I think if you listen to the conversations we have, the CEOs come on because they’re all competitive, they’re all type A. They like being challenged. They come here for a reason. We get almost more inbound requests to be on the show than we go out and ask. And meaningful, not just random cold PR pitches, but stuff we take, pitches we accept.

It is because all of those folks hear the show, they hear their peers on the show, and they want to beat them. I’ve asked about this a lot because in the early days I wanted to know why anyone comes on the show. I should understand it so that we can craft better and more compelling pitches for all the people we want to come talk to us. A lot of the answers I got were, “Because it is journalistically independent, the conflict is real.” 

That means, “Not only does the audience listen to it, but our own staff will listen to it.” If there’s some message that some CEO wants to get out, it is more likely that the staff will listen to it. Their employees will listen to it on Decoder rather than in an all-hands meeting. That is just the funniest thing in the entire world, but it’s also true.

I’ve heard a version of that many, many times. There’s pressure on The Verge, but we’re just old enough and we are famously spiky about things like being on background and off the record. Sometimes it comes to our younger reporters because they are testing to see if the younger reporters know. And then the younger reporters are like, “I’ll call Nilay,” and it goes away, which is fine. That’s my job in this newsroom. A lot of that has been offloaded onto the creator economy. It is not easy to pressure us, but it’s easy to go to some young creator who doesn’t even know that they shouldn’t be pressured or is excited about accepting conditions on access.

I see that left and right, all over the place. I don’t mean to call out anyone specific. I just think we play a role in the ecosystem where we are known for not doing that. I talk to young creators who are going to come up. Ben Smith was just on the show and he was noting that they’re all getting better at making journalism because the conflict is interesting. I see that ecosystem developing in parallel, but that’s where the pressure has really gone. It has not come to us very directly or it has stopped as we’ve gotten much older.

Questions or comments? Hit us up at decoder@theverge.com. We really do read every email!

Read the whole story
alvinashcraft
21 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Generative AI in the Real World: Local Voice AI with Pete Warden

1 Share

Pete Warden has spent his career on the frontier of small, local AI, first as one of deep learning’s earliest engineers (he coined the term “TinyML”) and now as founder of Useful Sensors and Moonshine AI, where he builds voice models that run entirely on-device. Pete joined Ben to make the case that local AI no longer has to be a compromise. They get into what it actually takes to run a capable model on a laptop today; why the voice interface’s bad reputation is a consequence of rough, early implementations rather than a reflection of current capabilities; and where he stands in the ongoing debate between general “end-to-end” models and the compound AI approach of chaining specialized models together. Pete also explains why he thinks browser-based inference could be an “iPhone moment” for local AI and why more and more enterprises are considering self-hosted local models over commercial options. “The shape of [LLMs] is perfect for running locally,” Pete says, and local models could be a boon to enterprises worried about cost, privacy, and stability.

About the Generative AI in the Real World podcast: In 2023, ChatGPT put AI on everyone’s agenda. In 2026, the challenge will be turning those agendas into reality. In Generative AI in the Real World, Ben Lorica interviews leaders who are building with AI. Learn from their experience to help put AI to work in your enterprise.

Check out other episodes of this podcast on the O’Reilly learning platform or follow us on YouTube, Spotify, Apple, or wherever you get your podcasts.

Takeaways

01.26 The usability gap is smaller than the marketing gap. The capabilities of local models are only a few months behind those from the big commercial companies, but because there’s no subscription revenue model behind local models, they often go unpromoted. “It’s very hard to make money off local models,” Pete explains, so the big companies aren’t focused on selling them. “Every company is going to go for the [product] that has an easy subscription revenue model. And that means you have a massive ton of marketing around all of these tools that are kind of like, ‘Oh, let’s have a little text box on a website.’ And so it means mostly that people have never heard of these local models.”

04.20 Local models are already good enough for most use cases. Pete compares the moment to the early web, when free alternatives like Apache eventually overtook expensive commercial servers. “All of these alternatives, once people actually had time to look around and they had a little bit of time to improve, they just wiped the floor with the commercial [offerings],” he points out. “I don’t know if we’re going to quite get there, but that’s the kind of pattern that I’m seeing.”

07.26 “The hardware barriers are a lot lower than people think.” Ben and Pete discuss what hardware you actually need to get up and running, from parameter counts, quantization (Q4, 8-bit), and VRAM requirements to the new Apple M5 Studio’s unified memory as a way to run very large models locally at usable speed. “The key thing is whether you can fit [your model] into your graphics card’s memory,” Pete says. “So with weight quantization, 9 billion [parameters] if it was 8 bits is like 9 GB. A lot of mid-end decent laptops that are shipping now have more than that.”

18.33 “It’s not that people don’t like voice interfaces. It’s that people don’t like bad voice interfaces.” We’ve solved most of the big problems, like dealing with background noise, phrasing, and speech in a range of accents—or at least have improved tools’ capabilities. However, “there’s no commercial incentive to kind of pull them all together,” Pete says. Most tools feel like they haven’t caught up to the LLM era, but “open source can be a really strong lever” to updating them, argues Pete.

28.26 We’re navigating the split between “LLM maximalist” end-to-end models (favored by big AI companies with the most capital) and the “compound AI” approach of chaining together specialized models from different sources. “If the future is end-to-end models, then only the people with the most money can actually build and train them,” Pete notes. Compound AI lets you “actually train all of the models independently” to accomplish your particular goals. While the performance of end-to-end models continues to improve, especially for multimodal models like Qwen or Gemma, using one can be a bit like choosing a Swiss Army knife over a tool specially designed to accomplish a single specific task, to use Pete’s metaphor. It may get the job done, but it’s probably not the most effective way to do it.

36.10 Voice capabilities in the browser could be a game changer. Embedding a model directly in the browser—Chrome has a built-in ~4B parameter model that’s accessible from any website via JavaScript, for instance—makes it part of the operating system. “Once you are able to transcribe fast and accurately in the browser, it’s a way for people to easily start experimenting with this stuff,” Pete explains. Could this be an iPhone moment for LLMs?

39:58 The “gravitational pull” is toward on-prem. Unlike most recent technological advances that depend on the cloud to function, LLMs are well-suited to running locally, even with no internet connectivity. Enterprises are grappling with concerns about cost, privacy, capabilities changing with no notice, or even the models they depend on disappearing. Hosting your own model, whether on your laptop or in your corporate infrastructure, gives you the stability to plan for the long term.

44:21 GPUs are fantastic for training but “complete overkill for inference.” Pete likens it to “trying to use an oil tanker to go and do your shopping.” Memory bandwidth is the real limiting factor, and it’s a problem that companies like Apple, with its new chip designs and unified memory bandwidth, are working on solving. “Even if you’re running on the CPU, if you have something that’s got high-enough bandwidth to pull 27 billion weights in a fraction of a second, then the rest of it is fairly easy in terms of actually doing the processing,” Pete says. “I think we’re going to see a lot of really imaginative solutions now that people understand what the workload looks like.”



Read the whole story
alvinashcraft
22 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

How Mbodi Is Solving Robotics’ Scaling Problem with Xavier Chi

1 Share

Robotics is having an AI boom, but don’t expect it to have a ChatGPT moment.

In this episode of Build Mode, host Isabelle Johannesen sits down with Xavier Chi, co-founder of Mbodi, a startup building AI software that lets people teach industrial robots new skills using natural language.

Xavier first joined Isabelle on the Startup Battlefield stage, where Mbodi became a crowd favorite with a live robot demo. Now, he’s back to talk about what happened after Battlefield and why he believes this generation of AI could finally help robotics companies overcome some of the problems that have historically made them so difficult to scale.

Xavier breaks down why traditional industrial automation still requires so much programming and customization, how Mbodi is working with ABB Robotics to bring its software into factories and warehouses, and why deploying robotics in the physical world creates a very different set of challenges from building traditional software. He also explains why robotics startups have struggled to build scalable software businesses and how generative AI could begin to change that.

They also get into what VCs are looking for in robotics startups, how founders can de-risk an investment before fundraising, whether the excitement around humanoid robots is justified, and why Xavier believes robotics won’t experience a single ChatGPT-like breakthrough. Plus, he explains why reliability is so critical on the factory floor — and how Mbodi recently achieved a 99.6% success rate during an eight-hour test.

They get into:

  • Why industrial automation is still surprisingly manual

  • How Mbodi lets people teach robots using natural language

  • Why labor shortages are driving demand for automation

  • How winning an ABB Robotics competition led to a major partnership

  • What Mbodi gained from Startup Battlefield

  • Why robotics software companies have historically struggled to scale

  • How generative AI could reduce the need for custom robotics integrations

  • What VCs want to see before investing in a robotics startup

  • Why robots don’t necessarily need to look human

  • Why software could capture more of the value in robotics as hardware gets cheaper

  • Why robotics won’t have a ChatGPT moment

  • Why reliability is so important when robots enter production

  • How Mbodi achieved a 99.6% success rate in an eight-hour test

Chapters:

00:00 — Mbodi’s unforgettable Startup Battlefield demo01:32 — Why industrial automation is still so difficult03:55 — What companies are actually using robots for06:33 — How Mbodi landed its partnership with ABB Robotics08:30 — Are startup competitions worth a founder’s time?10:00 — Why Xavier wanted to compete in Startup Battlefield11:18 — The hardest question Mbodi faced at Battlefield13:04 — Why robotics software has struggled to scale16:35 — Raising venture capital for a robotics startup18:02 — Convincing VCs you’re the right team19:12 — Do robots really need to look human?21:42 — What needs to improve across the robotics stack25:08 — What VCs want from robotics startups28:00 — How to de-risk a robotics investment30:06 — What’s next for Mbodi31:30 — When will robots enter our homes?35:35 — Why robots can’t afford to fail36:01 — Mbodi’s 99.6% reliability test

Subscribe to Build Mode on Apple Podcasts, Spotify, or wherever you like to listen. And watch the full videos on YouTube. New episodes of Build Mode drop every Thursday.

Hosted by Isabelle Johannesen. Produced and edited by Maggie Nye. Audience development led by Morgan Little. Special thanks to the Foundry and Cheddar video teams.






Download audio: https://www.podtrac.com/pts/redirect.mp3/traffic.megaphone.fm/TCML6393302995.mp3
Read the whole story
alvinashcraft
22 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

#562: DuckLake: The Lakehouse That's Just SQL and Parquet

1 Share
How many files does your query read before it reads any data? On some data lakes, you go through JSON and metadata files first, just to learn which Parquet files matter. DuckLake asks one SQL question instead. The metadata lives in a real database. The data stays in plain Parquet. That's the entire format.

Pedro Holanda joined DuckDB in 2018, when it was still a research prototype at CWI. He's the lead DuckLake developer. Guillermo Sanchez Dionis works on DuckLake and the new Quack protocol.

With Quack as the catalog, DuckLake handles 200 transactions a second under heavy contention. No other open table format comes close.

Episode sponsors

Six Feet Up
Talk Python Courses

Guests
Pedro Holanda: pedroholanda.org
Guillermo Sanchez: linkedin.com

PhD on progressive indexes: ir.cwi.nl
SQLite: www.sqlite.org
Litestream: litestream.io
boring hardware: talkpython.fm
DuckDB: duckdb.org
episode 491: talkpython.fm
Iceberg: iceberg.apache.org
manifesto: ducklake.select
DuckLake: ducklake.select
spec: ducklake.select
this diagram: blobs.talkpython.fm
Data inlining: ducklake.select
ducklake-dataframe: github.com
Polars course: training.talkpython.fm
CSV parser: duckdb.org
Zero-copy Arrow: duckdb.org
ART index: duckdb.org
async I/O: duckdb.org
v1.0: ducklake.select
Git-like branching: ducklake.select

Watch this episode on YouTube: youtube.com
Episode #562 deep-dive: talkpython.fm/562
Episode transcripts: talkpython.fm

Theme Song: Developer Rap
🥁 Served in a Flask 🎸: talkpython.fm/flasksong

---== Don't be a stranger ==---
YouTube: youtube.com/@talkpython

Bluesky: @talkpython.fm
Mastodon: @talkpython@fosstodon.org
X.com: @talkpython

Michael on Bluesky: @mkennedy.codes
Michael on Mastodon: @mkennedy@fosstodon.org
Michael on X.com: @mkennedy




Download audio: https://talkpython.fm/episodes/download/562/ducklake-the-lakehouse-thats-just-sql-and-parquet.mp3
Read the whole story
alvinashcraft
22 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

AI Model Month is Already Delivering Big Gains

1 Share
From: AIDailyBrief
Views: 67

September opens with a flood of releases as Gemini 3.8 Flash, Muse Spark 1.3, and ChatGPT Images 2.5 all land within days of Fable 5.1 and GPT-6 Astra. NLW breaks down where each one actually fits, why speed and cost efficiency now matter as much as raw capability, and what Meta's new Muse personal assistant signals about consumer agents. In the headlines: OpenAI's Navier-Stokes solution and the ugly credit fight around it, a class action over Claude Max usage limits, ElevenLabs eyeing an IPO, and Cognition's raise at $48 billion.

The AI Daily Brief helps you understand the most important news and discussions in AI.
Subscribe to the podcast version of The AI Daily Brief wherever you listen: https://pod.link/1680633614
Get it ad free at http://patreon.com/aidailybrief
Learn more about the show https://aidailybrief.ai/

Read the whole story
alvinashcraft
22 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet

1 Share
Macro view of overlapping textured paper sheets in white, pink, blue, orange and teal.

When Anthropic launched Claude Fable 5.1 this month, it centered the announcement around one benchmark result: its Terminal-Bench-Science score.

In this benchmark, a model gets a terminal and a real scientific research problem to solve independently. Fable 5.1 scores 52.6%, and Fable 5 scores 24.7%. By Anthropic’s scoring, the new model more than doubles the old one.

Anthropic’s published score was produced under conditions most users don’t have access to. The benchmark allows each model up to eight hours per task, and Anthropic has not said what harness or budget it used to get its numbers. When the benchmark’s leaderboard tested Fable 5, it ran the model through Claude Code at maximum effort and spent $14,180 across 210 attempts, about $67 each.

Most people don’t use Fable in a lab, so I wanted to know what an average user would get on these tasks. I recently tested Fable 5 and Fable 5.1 on everyday work and found them far closer than the benchmark suggests. That made me want to run the benchmark’s own tasks the way a home user would and see where the models actually differ.

The benchmark’s 70 tasks are public, so I pulled five of them, one from each science field, and ran both models myself. 

The tests

Terminal-Bench-Science has five categories, each with multiple tests. I chose one test per category that could run in a Python environment. Here’s what I picked:

  • Symbolic regression (mathematics) – A dataset with 100 variables and a hidden formula behind a yes-or-no label. The model must find a predictor that works on data it has never seen.
  • Lorenz-96 assimilation (Earth sciences) – Reconstruct a chaotic atmospheric model from a few uncalibrated sensors with unknown clock offsets. Grading is all-or-nothing on five criteria.
  • Reactor safety control (engineering) – Write a controller for a chemical reactor that finishes every batch as fast as possible without ever exceeding the temperature limit, across public and hidden fault scenarios.
  • Foraging cognitive model (life sciences) – Predict, trial by trial, which lever each of 20 mice will press, graded on sessions the model never saw.
  • Nanoindentation (physical sciences) – Extract material properties from raw indentation curves that include drift, adhesion, defects, and an unknown tip shape.

Each run got a plain terminal, and I set a $12 limit and 60 turns for each test. The full set of ten runs took about 12 hours.

Symbolic regression

This was the only test where a model passed the benchmark’s hidden test. Fable 5.1 worked for 27 turns, found the hidden structure, wrote a predictor, and stopped on its own after 11.8 minutes, 27,088 output tokens, and cost $1.96 to pass this one test.

Fable 5 used all 60 turns over 53.5 minutes, generated 39,461 output tokens, cost $4.20, and failed. I ran Fable 5 a second time to rule out bad luck. It used all 60 turns again, took 60 minutes, generated 60,608 output tokens, cost $6.38, and failed again.

Lorenz-96 assimilation

This was the most expensive pair of runs. Fable 5 hit the $12 cost limit at 45 turns after 97.7 minutes and 92,091 output tokens, ending at $12.63. Fable 5.1 used all 60 turns over 126 minutes, generated 89,789 output tokens, and cost $10.70. On the public leaderboard, Earth sciences is also the field where Fable 5 scores close to zero, and both models failed it here.

Reactor safety control

Neither wrote a controller that passed the grader’s scenarios. This run produced the most output tokens, 157,710 generated by Fable 5.1. It hit the 60-turn limit after 40.9 minutes and cost $11.53. Fable 5 hit the cost limit at 49 turns after 63.8 minutes, 121,978 output tokens, and $12.04. 

Foraging cognitive model

This was the longest run of the testing series. Fable 5.1 was the only model that declared itself finished. It built a model, tested it against its own scoring loop, and declared it done at 43 turns after 53.5 minutes. It created 65,518 output tokens and cost $5.65. But the official grader rejected it. 

Fable 5 never declared anything. It hit the $12 limit at 60 turns, after 139.3 minutes and 62,587 output tokens, ending at $12.13. 

Nanoindentation

Both failed. Both spent most of the run reading raw curves and writing code to segment them. Neither produced a results file the grader accepted. Fable 5.1 ran out of turns at 29.6 minutes, 115,687 output tokens, and $10.91. Fable 5 ran out of money at 48 turns after 34.1 minutes and 114,239 output tokens, ending at $12.59. 

Results

Here are the results by the numbers.

MetricFable 5 scoreFable 5.1 score
Tasks solved0 of 51 of 5
Output tokens430,356455,792
Total cost$53.59$40.75
Total time388 min262 min
Runs ended by cost limit40

The benchmark scores models on all 70 tasks with three trials each, and Anthropic’s 24.7% and 52.6% come from that full suite. The independent leaderboard puts Fable 5 at 21.4%, close to Anthropic’s figure. Fable 5.1 is not on the independent leaderboard yet, so its 52.6% is Anthropic’s number alone. My results, 0% and 20%, are below both. Five tasks are a small sample. 

Getting these results by chance is plausible even if the published scores are exactly right, so this run neither confirms nor contradicts the doubling claim. The direction matched, since the new model did better. The one task Fable 5.1 solved was in mathematics, which is also the field where the leaderboard shows Fable 5 performing best.

What I think

I don’t think a regular user will see much difference between Fable 5 and Fable 5.1. I only ran a small sample of tests, so I can’t prove or reject Anthropic’s benchmark results. But what I saw suggests that gap won’t reach the average user. The one difference that did show up was on the bill. Fable 5.1 failed faster and cheaper, and it never hit my cost limit, whereas Fable 5 hit it four times.

Suppose your work looks more like the benchmark tasks; the harness and the budget matter as much as the model. With a purpose-built harness, hours per task, and a much bigger budget, you may get closer to Anthropic’s numbers.

The post Fable 5.1 vs. Fable 5: Results on a real-world budget, not the spec sheet appeared first on The New Stack.

Read the whole story
alvinashcraft
22 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories