Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
162423 stories
·
33 followers

What AI gets wrong and what failure teaches us

1 Share

Jennifer Neville is a partner research manager at Microsoft who’s built a career around understanding and advancing AI for real-world use, and much like the human-AI interactions she’s been studying, her early-career path was multiturn: math, then physics; cognitive science, then work; and finally computer science—despite her best efforts to avoid the field. 

In this conversation with Principal Applied Scientist Chad Atalla, she explores the role evaluation plays in pushing the performance boundaries of today’s AI systems to meet user needs and the “surprising failures” that emerge when models are tested beyond traditional benchmarks. Neville also shares practical guidance for working with current AI systems and discusses why looking closely at data matters when results defy expectations, and what decades of AI progress have taught her about predicting what comes next. 

From an unexpected career trajectory to the frontier of AI interaction and learning, this episode asks a larger question: what can we learn when the path—whether human or artificial—doesn’t unfold the way we expect? 

Transcript 

[MUSIC] 

JENNIFER NEVILLE: There were times that I felt like I was spinning my wheels. Things weren’t working. I would get, kind of, dejected. And then we’d get to a point where we did learn something, and it was just the, the emotional thrill of that was, that was really what hooked me. The being able to, kind of, understand something that no one else understood yet …

CHAD ATALLA: Sure … 

NEVILLE: … because it’s just at the frontier of what we know about things. 

STANDARD INTRODUCTION: This is the Microsoft Research Podcast, where Microsoft researchers—driving advancement through fundamental science and technology research—explore the who, how, and what’s next in computing and AI. 

[MUSIC ENDS] 

CHAD ATALLA: Hello, and welcome. I’m Chad Atalla, an applied scientist here at Microsoft Research. 

Today, I’m joined by Jennifer Neville, a partner research manager at Microsoft Research and the Samuel D. Conte chair professor of computer science and statistics at Purdue.

Her research examines machine learning and AI for interactive domains and structured data, looking at how the data points that AI systems are trained on affect their behaviors and how that aligns with what users actually want.

Across that research, she has published more than 130 papers with over 10,000 citations and received honors such as a National Science Foundation CAREER Award, a spot on IEEE’s “10 to Watch” in AI list (opens in new tab), and best paper awards from the International Conference on Data Mining (opens in new tab) and the International Conference on Learning Representations. 

What amazes me most about Jen’s work is her foresight and ability to have a through line that stays consistent as the field of AI develops, and I’m excited to hear more about how she got to where she is today. 

Jen, thank you for joining me. 

JENNIFER NEVILLE: Thanks for having me. 

ATALLA: Awesome. Well, I would love to learn about how you got here today. And let’s rewind way back. When I was a kid, I wanted to be an astrophysicist, but of course, here I am with a computer science background. When did you know that you wanted to go into computer science? 

NEVILLE: That’s a good question. I … when I was growing up, I wanted to do anything but computer science because that’s what my dad was in, and I wanted to do anything but what my dad did. So when I went to college, first I majored in math and then in physics, but I really just couldn’t vibe with those majors. 

So I dropped out of college for a while, and when I went back to college again, I majored in cognitive science instead, which also was too squishy for me. Didn’t have enough math in it. At that point, if somebody had told me, “AI is the thing you should be doing because it combines the cognitive science with the math and computational thinking,” I think I would have saved myself a lot of time, [LAUGHTER] but that didn’t happen. 

And so I worked for a while. And then when I went back to school, I decided to major in computer science because what I wanted to do was think about working with data and dealing with data. And that’s when I found AI. Just by happenstance. I didn’t even really realize that it was part of computer science. 

ATALLA: Awesome. Well, it’s one thing to go into computer science and want to work on data or AI and another thing to want to be involved in research on that front. So what sorts of questions or big ideas sparked your research drive? 

NEVILLE: Yeah, that’s actually an interesting story, as well. I only got into research because I was in the honors program in my computer science degree, and as part of the honors program, you have to do a research project [LAUGHTER]. 

ATALLA: It’s mandatory, yeah. 

NEVILLE: It’s mandatory. And so when I talked to professors about what the research project should be, they actually said, “Well, you have to decide what topic you want to work on.” 

And at that point, I was interested in data and I was interested in AI, and I thought a lot. I did a lot of reading. And what I decided that I wanted to investigate was how to do data mining on web data that was interconnected, and so that … when I decided that was the question I wanted to work on, I got pointed to a particular faculty member … 

ATALLA: Nice. 

NEVILLE: … who had just started working in this nascent field at the time that was called statistical relational learning. And we did my project in that, I published a paper at a workshop, and I was, kind of, hooked. 

ATALLA: OK, yeah. 

NEVILLE: So I hadn’t planned to go on to grad school, but that experience … 

ATALLA: Yeah. 

NEVILLE: … made me want to go on to grad school and continue. 

ATALLA: Nice. What part of it do you think hooked you? Was it, like, the thrill of doing the research? Was it the environment of the conference and what academic publishing looks like? 

NEVILLE: It was really the thrill and the process of doing research. So it was a long project. There were times that I felt like I was spinning my wheels. Things weren’t working. I would get, kind of, dejected. But my adviser would be, kind of, like, you know, supportive and positive, saying, “No, keep going. We’re learning something.” And then, and then we’d get to a point where we did learn something, and it was just the emotional thrill of that was, that was really what hooked me. The being able to, kind of, understand something that no one else understood yet … 

ATALLA: Sure. 

NEVILLE: … because it’s just at the frontier of what we know about things was, uh, it’s just, it’s like a drug almost, right? [LAUGHTER] So it’s kept me in research for this long, that same, that same … chasing after that same feeling. 

ATALLA: Gotcha. Well, love it. Yeah, you joined Microsoft in 2021 from academia. And in fact, you’re still a professor, and you’ve been teaching and advising for 20 years. What motivated the shift to part of your professional life being research in industry, and how would you characterize the difference between industry research and academia research? 

NEVILLE: My research, kind of, spans the spectrum from theory to application. When I did my first sabbatical after I got tenure, a lot of colleagues that I had at the … that were at the same point in their career went off into industry labs for their sabbatical and never came back to Purdue. I thought about which direction I’d want to go, either more theory or more applied, and I thought at that point in my career, it’d be better to explore the theory side because then I might actually go back to Purdue. 

And so I went to the Simons Institute in Berkeley for my sabbatical, and that was very theoretical. It was a great experience, but it was kind of seamless to transition back to academia. On my second sabbatical, I went the other way, which was to come to MSR [Microsoft Research] and do a sabbatical here. And of course, being able to see how algorithms in theory, kind of, hit the—where the rubber hits the road with respect to how they behave in practice and real systems and with real users and real data, that’s hard to turn back from. And so that’s now why I’m still here. 

ATALLA: Gotcha. 

NEVILLE: So I think the difference between research in academia and industry depends on … really is affected by the target of where, where you’re aiming the research. And so I think fundamentally, it feels very similar, the questions you would ask in academia and industry, but in industry, your ability to apply it at scale in real systems on real data is really very different from academia, and in academia, I think you end up asking questions that are more abstractions that cover applications across a lot of different domains at once. And that’s how you get funding from places like DARPA [Defense Advanced Research Projects Agency] and NSF [National Science Foundation]. 

But in industry, it’s a little easier to just, you know, actually get your hands dirty and, and do it in the real systems. And then your, kind of, target is the products and the company’s interests. And so as … that sort of changes, maybe the types of questions that you would ask or investigate. So I think it’s actually great to be … 

ATALLA: Yeah. 

NEVILLE: … in both places, right? Like, there’s lots of synergies across the two, and I think there’s lots of opportunities to go to industry and come back to academia or to go start in industry and go to academia. But right now, if you’re doing work on AI systems, I think that industry is really the place to be. 

ATALLA: Gotcha. Yeah. Perhaps useful advice to folks who are deciding which way to take that decision in their career right now. 

I’ve been here at Microsoft for a little over six years and have been working on AI evaluation for much of that time, and there are, of course, a number of daunting fundamental questions and challenges in the space. 

And I initially came across your work and became aware of your work because part of your work proposes some solutions to some of these problems and helps to point out these problems. And my immediate reaction was relief that we have someone wonderful like you working on these things, and you’re going to take care of it for us. And I really appreciate how you balance critiquing AI evaluation and also providing solutions along the way. 

But I understand that that’s just a small sliver of your broader research agenda. And so how would you describe the current research charter for the team that you’re leading here? 

NEVILLE: Our team is called the AI Interaction and Learning team. What we really focus on is trying to push the frontier of behavior of these AI systems in realistic work environments. And so what that entails is studying where’s the boundary of the performance of the current systems, how do users experience that in practice with real workflows, and then how do we improve that and push the performance of the models for these kind of tasks—complex tasks—that users are working on. 

The reason we focus on evaluation is that we find that the standard way that ML and AI people evaluate are through benchmarks that are fairly simple with respect to how people would actually use them in practice. And so we find that the first thing that you really need to do is ask, what do we want out of these systems? And design practical evaluations to put the systems in those kind of environments. So that means things like multiturn behavior, collaborative environments, long-horizon tasks that users are doing. 

And then once we can see where the performance gaps or problems that happen in those evaluations [are], then that also gives us the knowledge of where, sort of, theoretically or algorithmically do we need to improve these systems. 

So we start with evaluation, but ultimately, what we’re trying to do is develop better algorithms, models, estimation methods … 

ATALLA: Sure. 

NEVILLE: … to push the performance of the models. 

ATALLA: Yeah, it’s a gateway to understanding and therefore being able to push the boundary and improve these, these systems. 

NEVILLE: Yes. 

ATALLA: So you mentioned multiturn interactions and collaborative environments. You have a couple of recent papers that examine these things specifically, like how LLMs may get lost in multiturn conversations or how they may corrupt documents in these, sort of, agentic knowledge work scenarios. Can you tell me a little bit more about those research projects specifically? 

NEVILLE: Sure. We … those are examples of cases where we had to, sort of, have innovative insights as to how to set up the evaluation to really be able to control and test the model’s behavior in those scenarios. 

So for multiturn conversations, what we did is we took single-turn benchmarks—maybe I should back up and say one of the things we observe about users in practice is that they often underspecify what they’re trying to do. They don’t know how to give a complete specification to begin with. We sometimes forget this as computer scientists and software engineers, but lay users might not be able to fully specify things in the first turn. And so they, kind of, figure out what they’re doing over multiple turns and, sort of, add conditions and clarifications throughout the turns. 

So what we did was we took single-turn benchmark datasets that are public, and we changed the fully specified single complex turn of instructions and had the, had models that simulated users providing that clarification over multiple turns and then studied how do the models—how do all of the current models that we have—how do they behave when you get this kind of task specified over multiple turns. 

And we found that the performance that we see in single-turn scenarios—which is very high because, of course, we’re optimizing to that as model builders—actually degrades significantly over multiple turns. And so that was something that was a fairly surprising finding because nobody had evaluated in that, … 

ATALLA: Sure. 

NEVILLE: … in that way before. But when we released that research into the world, the users of the systems all resonated with, … 

ATALLA: Sure … yeah. 

NEVILLE: … “Well, that’s been my experience. I knew that that was happening. How do, how do we fix that?” And one of the … so we’re working on reinforcement learning methods to actually help the models learn better in these environments to, kind of, fix that behavior. 

But a practical solution that we talked about in the paper is if you end up having the model, kind of, get confused—if you specified something over multiple turns—just, like, stop the chat, erase everything, go back, and now, given what you know, give it a fully specified single turn, … 

ATALLA: I see, yeah. 

NEVILLE: … and then it will behave better in that kind of interaction. 

So maybe that points out that although fundamentally we’re looking for algorithmic improvements to the models, sometimes in these projects, we end up coming up with insights of how users could change their behavior … 

ATALLA: Sure. Yeah. 

NEVILLE: … to get better outcomes from the models, like, where they are right now … 

ATALLA: Yeah … 

NEVILLE: … with their performance. 

ATALLA: Yeah. Well, we noted that one of the themes of your research is how we can improve these systems to bring them better into alignment with what users actually want. And you mentioned a case here of, “Hey, well, we know that users realistically are often not fully specifying what they want in their first turn, and we do have these long multiturn interactions.” And after you put this paper out, they expressed resonance with this frustration about getting lost in multiturn conversations. 

How do you think about understanding the users’ wants and needs and experience, and how does that relate back to what you do with, as you said, simulating a user, for example, to run these longer multiturn evaluations? 

NEVILLE: Yeah, that’s … getting to see at scale what users are trying to do and what kind of failures and successes they have in the systems is one of the advantages I think you get if you’re in an industry lab and you get to see data at scale like that. 

So we have done a lot of work with the product groups here internally at Microsoft analyzing consumer logs at scale to figure out what are the main patterns of failures that, that users are experiencing. And that is the … provides the kind of insight or motivation for a lot of the, a lot of the work that we do. And there’s a lot of great product groups that are also analyzing the successes that users have, which would allow you to then even recommend to users how … what are the kind of tasks that they’re going to be able to do successfully in the systems that we have now. 

I think if you think back when search engines first started, people had to learn how to interact with search engines … 

ATALLA: Sure. 

NEVILLE: … in a way to effectively get the information that they’re looking for. We might have forgotten that that happened now because it was so much a part of what everybody would do before all these AI systems came out. But now I think there’s a transition away from the search engines to these chat-based systems, and I think there’s going to have to be some learning from the users, as well, … 

ATALLA: Right, yeah. 

NEVILLE: … in this new environment. And I think analyzing the things that are successful in terms of what users are doing is a great way to start recommending … 

ATALLA: Yeah. 

NEVILLE: … and tutoring or teaching the users in the same way. 

ATALLA: I love how you’re pointing out the utility of looking at real data, but what does that actually mean? Are you genuinely looking at real data? How is privacy coming in here? What sort of processes do you have for responsibly leveraging data in the wild? 

NEVILLE: Oh, yeah, that’s a good point. Thanks for bringing it up. We’re not actually looking at the data. We are analyzing the data at scale to extract patterns of failures or successes in a privacy-preserving kind of way. There’s also a lot of, a lot of restrictions that we can’t actually see the data eyes on. We can only, sort of, push methods to process the data in very restricted environments. And then, um, and then these, sort of, higher-level insights that are privacy preserved are the things that then fuel, like, what we would do algorithmically later on. 

So just to be clear, we’re not fine-tuning on your data or your responses, but the more details you can give in your responses, whether it’s through donated information, when you give a thumbs down and you give a description of it, or if it’s in the, sort of, chat interaction you have with the models, those can eventually get distilled into things that are going to improve the models not through people actually looking at what you’ve been doing. 

ATALLA: And you noted this shift from a search being the dominant modality to these multiturn chat interactions. But now we’re also seeing agentic and, you know, collaborative knowledge-work-style tasks becoming even more relevant and noted that you did some work on that front, as well. Any other interesting takeaways that you’d love to share with our listeners? 

NEVILLE: Yeah. So we have both some theory work that tries to characterize the types of tasks that are fundamentally hard for transformers to compute inside the model as well as more simulation-based work, where we look at complex tasks that are conducted over long-horizon workflows, where you take, for example, documents and you do repeated edits over those documents. 

In those environments, the complexity of the task for the models or even the agents is to track not only what the user’s intending to do over this long-horizon work activity but also understanding what’s the current state, what information is relevant, what is not relevant, how things have changed. And those are things that agents are starting to be fairly good at, but there’s particular kinds of tasks that are easier than others. 

ATALLA: Sure. 

NEVILLE: And so something that we look at is what are the types of tasks where complexities become too much for the current agents and how to solve that. What is the best way for the tool usage with the agents? Is it better to bring a human in the loop? How can we know that things have gone awry? So having … even having the agents, kind of, monitor themselves and assess whether they answered something correctly or they should, kind of, roll back to a previous state. I think those are all open questions right now. 

But we have … we do have some work showing that again we get surprising failures as we get longer and longer into these workflows because errors, if you don’t catch them, can accumulate … 

ATALLA: Yeah. 

NEVILLE: … and start to confuse the AI even more as they, as they have to carry out the task over repeated interactions. 

ATALLA: Yeah, that phrase “surprising failures” is interesting to me. It implies an expectation of, “Hey, this should do better, and I’m surprised that it failed in this specific way.” Do you feel that you see more of these surprising errors, or are a lot of the errors, like, expected, you would understand that it would fail in, in this way or that? 

NEVILLE: Yeah, I think there, I think there definitely are expected failures because those … well, there’s fewer expected failures, of course, as we work to make the systems better and better. 

ATALLA: Right. If we can expect them, then we can fix them. 

NEVILLE: That’s right. But things like hallucinations are the types of failures that have been characterized for a long time in the community and that there are specific benchmarks and methods to try to reduce those things. 

I think that my team ends up looking for things that, that show up as almost surprising failures because they’re not the traditional failures and they’re often hard to isolate and show that they’re happening because they’re very subtle and don’t show up as just, kind of, [an] incorrect answer to a math problem or a hallucinated case in a law, you know, a legal opinion but are this sort of subtle loss of semantic content in the documents. And I think that if I look back on what we’ve, sort of, investigated with the team, we have a kind of hypothesis that, that we think as humans, the things that are hard for us to do are going to be the things that are hard for the model to do. And in fact, that’s not often the case because we have to think about what’s hard for the transformers actually to compute, what’s hard for their retrieval systems to retrieve … 

ATALLA: Right. 

NEVILLE: … from the underlying content. 

And so I think where the “surprising” comes from is things that we might think as humans are very simple to do or that we think if there was a task previously that was done correctly, we think reliably now the same task again should be done correctly. But in fact, that’s not necessarily the case with LLMs. And so the surprising, kind of, creeps in, I think, with our own expectations based on human intelligence about what’s hard or what’s easy. 

ATALLA: Yeah. And so between that confusion that may exist for users based on what they expect these systems to be good on, perhaps based on reflecting on human intelligence, and all of the hype and marketing that exists out there, what do you want real users to take away from this? Or how would you caution them to think about the capabilities and expectations that they may have for AI systems given your work here? How can they kind of cut through that noise and build better expectations? 

NEVILLE: Oh, that’s a good question. I think that, I think the current AI systems have a lot of capabilities that are going to be very useful to people in work environments right now. I think that the takeaway from our work is that you shouldn’t expect them to be 100% successful across the board. You should be checking the answers that you get back. And if you get back something that you think is incorrect, you shouldn’t necessarily assume that the model can’t do that at all. But you should think about asking again or asking it in a different way, and you might get a better answer the next time. 

That is hard to work into the way that we work right now because I think that you can’t … they’re not ready for things to be 100% delegated to them. But if you can figure out how to use them in a workflow with you, with oversight and verification, I think they can be very useful to do things, improve, you know, productivity and the speed with which you get things done. 

So … and maybe the other thing to say is, you know, sort of bear with us because [LAUGHTER] we’re developing these systems in real time as they’re being released. And so your use helps us actually push the boundary of the models because it gives us examples of things that can and can’t be done. 

Maybe another thing I would like to say is that something that maybe users don’t understand is that in the past, with search-based systems or recommender systems, the kind of feedback that we were able to give as users was just, kind of, like thumbs-up, thumbs-down. But now we’re in an environment where actually we could give feedback with much higher fidelity that can be super useful to improving the models downstream. 

So if you’re working in, kind of, a chat environment with a model, understand that if something has gone badly, it’s actually helpful to say how it’s gone badly, to … and in your textual interaction or speech interaction to really convey, like, exactly what went wrong and what you expected and what you got instead. Because that actually will be used as we update the models. And so you can think of that as users, your role can be to train the models to do the things that you would have wanted it to do, but it couldn’t do right now. 

ATALLA: Awesome. Interesting call to action there. 

And I’m just curious. You talked about the transformer paradigm that we’re in now for these sorts of large language systems. Where do you see the science of large AI systems going next? Broadly, maybe that’s a really hard question, but at least in your interest and research directions, where do you see the science of these systems going next? 

NEVILLE: That’s, that’s a, that’s a big question. [LAUGHTER] I think we’re … the research community is simultaneously trying to figure out if the transformer architecture is the right underlying model to use because there are certain known limitations of how computation can work in transformers. A lot of those limitations can be solved by the wrapper around the models—the agentic harnesses that we have around the models right now—and the reasoning that goes on, on the back end. 

So I think an open question is, how much can we deal with the limitations of the transformer architecture with this, sort of, wrapper foundation around it and what needs to be addressed with, sort of, a fundamentally different … 

ATALLA: Sure. 

NEVILLE: … architecture underlying things? I think that we’re going to see rapid development of, sort of, alternative architectures and models as well as elaborate wrappers around the current models that we have. 

I think that we will see in the next generation of work that we will have to focus on longer-term interactions with the models and having the models understand the world within which they’re working in a much more tangible way than they’re doing right now. But I think I’m very bullish on the research that’s happening. I think we will, you know, be successful at this. 

Maybe going back to when I started, my first internship was at AT&T Labs back in the year 2000. And I remember very clearly sitting at lunch in my internship playing go and talking about how we could get, you know, AI models to be able to, you know, play this game and how it was harder than chess and what could we do. 

And, and I was learning how to play go at that point. And I was learning on a 9-by-9 board. I don’t know if you’ve ever learned how to play go, … 

ATALLA: I have not. 

NEVILLE: … but you don’t learn how to play on the big board … 

ATALLA: Wow. 

NEVILLE: … because it’s, it’s too hard even for humans, right. So you start on this 9-by-9 board. 

And I was, kind of, watching myself learn how to play and think about how the algorithms would learn how to play. And while we were doing that, we were, you know, sort of, conjecturing about how long it was going to take AI to be able to do this. And so it’s funny looking back because, you know, at AT&T Labs, a lot of the, sort of, luminaries in, you know, machine learning and AI research were there. People were saying, “Oh, it’s going to take 50 years to be able to do this.” You know, some people say 100 years. But, of course, now that’s already a solved problem (opens in new tab). 

ATALLA: Yeah. 

NEVILLE: And so looking forward, there are things that I might think as a researcher, because it’s so hard for us to get past this sort of boundaries we see right now in behavior, there’s a tendency to think, “Oh, it’s going to take 10 years,” or “It’s going to take 25 years.” 

But I think given the pace of research and the number of people researching, you know, doing research in this field and the scale at which these models are running, I think it’s actually going to happen much more quickly than I would have thought as that person, you know, that new AI student in the year 2000. So … 

ATALLA: Wow. Well, in the spirit of that, looking back to the year 2000, let’s imagine now jumping 20 years forward into the future or just 10 because things are moving so fast. 

NEVILLE: [LAUGHS] Yeah. 

ATALLA: What mark would you like to have left on research in this space? 

NEVILLE: I think the, the things that have really been the, sort of, common thread through my research career is thinking about complex systems and how to get AI and machine learning to work in these complex systems. So—and even right now, that’s what we focus on with collaboration, and we’re thinking about multi-human, multi-AI interaction. I think … if I think really in terms of sci-fi kind of outcomes, I think we’d be successful if in the work environment, we start to see organizations of both humans and AI working together on teams to accomplish things and innovate, and if we can get that working correctly, then I will feel like my research has been successful. 

ATALLA: Awesome. Well, as we wrap up here, I thought it could be fun to do a lightning round, which means quick questions, quick low-stakes answers, just whatever first comes to your mind. Sound good? 

NEVILLE: Sure. 

ATALLA: Awesome. What’s one piece of advice that stuck with you or changed your perspective? 

NEVILLE: Uh, OK. You said short answers, though. [LAUGHS] 

When I first started in computer science, I, uh, actually, Doina Precup (opens in new tab) was the TA in my first computer science class, and she said to me when I was, sort of, being overwhelmed and feeling like I didn’t … everybody else knew everything and I didn’t know anything, she said, you know, don’t, don’t be afraid to ask questions because people might appear that they know what’s going on, but in fact, they don’t know what’s going on, and they will actually really appreciate that you’re courageous enough to ask a question because they’re probably thinking the same thing. And, and it would help both you and them if you are, if you will ask the question. 

And so I’ve taken that to heart over the course of my career. And it’s very hard to, you know, you can be afraid and feel like you’re going to appear stupid to ask the question. But undoubtedly, along the way, whenever I had these feelings of fear and … but had the courage anyways to ask the questions, I did find that people were really appreciative. And, you know, people would say afterwards, “Oh, I was thinking the same thing. I’m so happy that you asked that question.” So, so be courageous and ask the question. 

ATALLA: Awesome. What’s a lesson that you learned the hard way? 

NEVILLE: Look at the data. [LAUGHTER] Look at the data. 

So in machine learning we have a tendency—this applies to the benchmarks. We have, we have a test set. We run our model. We learn it. We apply it to this test set. We generally get a number that’s our metric that we’re trying to push. And so we think if the number is going higher, we’re doing good. And if the number doesn’t move, we think, “Oh, well, then the problem is too hard.” 

But I have learned the hard way multiple times that when things are not working out, if you take the time to look at the data, you might actually find that there’s errors in the data. The data is not what you expected. The data is different from the distribution that you train the model on. And so anytime … I guess maybe that makes me a data person. [LAUGHTER] But often when things are not behaving the way that we expect, we go and look at the data to try to understand. 

ATALLA: Yeah. Makes sense. It’s a good one. It’s a hard one. 

NEVILLE: [LAUGHS] Hard to remember. 

ATALLA: Yes. 

NEVILLE: And that … maybe that’s an important thing is, even though I’ve learned that lesson, … 

ATALLA: Sure. 

NEVILLE: … I’ve had to re-learn that lesson at least half a dozen times throughout my career. 

ATALLA: Yeah. Well, on the flip side, what’s an accomplishment that you’re most proud of? 

NEVILLE: I think I would say the fact that I was able to get tenure and raise a toddler at the same time. 

ATALLA: Wow, yeah. 

NEVILLE: I had my son right before we started—my husband’s also faculty—so before we started our faculty positions. And so I think the fact that I was able to get tenure, my son was healthy and happy, and I managed to stay married [LAUGHS] … 

ATALLA: Yeah, indeed. 

NEVILLE: … is probably the biggest accomplishment that, uh, yeah … 

ATALLA: Big accomplishment, yeah. And last question here. What’s one way that art or nature has influenced your work? 

NEVILLE: Hmm, I’m not sure if this would really count as nature, but one thing that I’ve always been observing as an AI person is how kids, people, animals, insects seem to learn about the world and the environment and change their behavior. And so I fully admit that I was very nerdy as a, as a parent with a young child. I would take, you know, ideas that we have from machine learning and try to see if I could help my son learn better. 

For example, if he’s learning to try to, like, flip over when he was an infant, I’d say, I’d say, “OK, let me give you a positive training example. [LAUGHTER] Like, here, take your legs and do this,” and … 

ATALLA: Yeah, yeah. 

NEVILLE: … and then see the effect on, you know, how he learned but also then watching him learn and, and then thinking about how that, you know, might be worked into the algorithms and models that we developed. That’s probably the biggest interaction that I’ve, like, sort of continually returned to over the course of my career. 

ATALLA: That’s beautiful. Yeah. 

NEVILLE: Thanks. 

ATALLA: Well, Jen, thank you for sharing your story and insights. It’s been a pleasure having you here. 

And to our audience, thank you for tuning in. If you would like to learn more about the work of my colleagues here at Microsoft, check out the Microsoft Research page at aka.ms/research or [MUSIC] check out the other episodes of this podcast. Thanks. 

[MUSIC ENDS] 

Opens in a new tab

The post What AI gets wrong and what failure teaches us appeared first on Microsoft Research.

Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

The results of the 2026 Developer Survey are here!

1 Share
Below, we’ll highlight some of the results we found interesting about what's going in the life of technologists, their technologies, AI usage, and more.
Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Tales from the 2026 Developer Survey results

1 Share
Ryan chats with Erin Yepis, Senior Analyst at Stack Overflow, about the results from this year’s Annual Developer Survey, including the overwhelming daily usage of AI coding assistants despite lingering developer trust issues, the critical role of well-organized documentation in providing verifiable context to mitigate AI hallucinations, and the evolving ways developers are shifting away from active community posting in favor of passive knowledge consumption.
Read the whole story
alvinashcraft
1 minute ago
reply
Pennsylvania, USA
Share this story
Delete

Radar Trends to Watch: October 2026

1 Share

In addition to the nearly constant stream of model releases, in September we’ve seen price drops, new kinds of models, proofs of long-standing problems in mathematics, and continued investigations into models escaping their sandboxes. (Axios reports investigations into over 10,000 incidents.) AI has infinite patience and is fundamentally probabilistic. Given a difficult or impossible task and an unlimited token budget, an agent will eventually attempt to solve the problem in ways that you don’t expect, and may not want. It’s easy (and correct) to blame inadequate security procedures at the frontier AI labs, but AI adopters must be careful not to make the same mistakes. The humans using AI need to be accountable for what their agents do.

AI Models

Model choice is starting to hinge on price and specialization as much as raw benchmark leadership. Alongside general chat models, there are now decision models that never chat, spatial models built for robot planning and camera control, forecasting models sized for a single task, and cybersecurity-specialized models kept behind an invite-only program. Specialization leads to greater efficiency and lower costs, at least in the short term. In the long term, specialized models may succumb to the “bitter lesson.”

  • Anthropic has released Claude Opus 5.5, which it claims has performance similar to Fable 5.1, and hence similar restrictions. It’s faster and requires fewer resources to run. Anthropic has dropped prices 20% for input and output tokens and 60% for cached reads. Not to be outdone, OpenAI released GPT-6 Sol and Luna, with 50% price reductions.
  • Anthropic has also announced Fable 5.1 and Mythos 5.1. The most significant change appears to be a 75% price reduction for cache reads, which might translate into significant savings for long-running jobs; Anthropic estimates 25%. Mythos is only available to trusted partners. Simon Willison used Fable 5.1 to animate his pelican-riding-a-bicycle pseudobenchmark.
  • Anthropic released Sonnet 5.5 with claims that it’s 30% faster and 30% less expensive for most work. The new model has security limitations similar to those applied to Opus and Fable; it routes to Sonnet 5 if it’s asked to do anything out of bounds.
  • And finally, as September closes, Anthropic announces a marketplace for Claude plugins and connectors. At its launch, Claude Marketplace had over 2,000 items.
  • OpenAI has released GPT-6 Astra, with claims that the company has achieved AGI (artificial general intelligence). Astra’s excellent benchmark scores appear to depend on the use of an unreleased harness. OpenAI has also released GPT-6.1 Sol, with per-token price reductions and claims that it is close to GPT-6 Astra in capabilities
  • OpenAI has solved the Navier-Stokes existence and smoothness problem, a mathematical problem in fluid mechanics. This development raises an ethical question: Did OpenAI train its system on the work of two mathematicians who were close to solving the problem themselves? It also raises practical questions about the future of mathematics. Decorated mathematician Terence Tao asks whether “the collection of good, fruitful open problems is now being mined in a non-renewable fashion.” An advisory group has been formed to help OpenAI make decisions about releasing mathematical results.
  • Google has released Gemini 3.8 Flash TTS and Flash-Lite TTS. Voice options aren’t limited to a prebuilt library. These models have APIs that allow developers to describe the voice that they want or upload a sample. These custom voices are then assigned an ID so they can be reused.
  • Google has announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking. These models are designed for live, near-real-time conversation. They can process video input. Live Extended Thinking can reason and speak at the same time.
  • Google has released Gemini 3.8 Flash and Flash Cyber. Flash appears to be similar to frontier models on most benchmarks, with Computer Use being the biggest exception. Flash Cyber is specialized for vulnerability detection and mitigation and is only available to defenders in the Fairwind Program.
  • Google has also released TimesFM-3, a small specialized model for multivariate time series forecasting. The model weights are available on Hugging Face with a license that only allows for noncommercial use. (Source code is available under the Apache open source license.)
  • Xiaomi released MiMo-V2.6, its latest large language model. It’s fully open sourced, and based on benchmark results, Xiaomi claims that MiMo is the strongest open model to date. What’s more interesting is the claim that MiMo only cost $3.5 million to train.
  • TypeSafe’s new decision model, Jev, is unlike anything else we’ve seen. It doesn’t chat; its output is always strictly typed and accompanied by probabilities that estimate correctness. It’s much faster and less expensive than other leading models. It isn’t open source, but there are already many open source clones.
  • Ollaya is similar to Ollama, but for running decision open models like Laya (a clone of Jev) locally.
  • World Labs has released Atlas, a model for “spatial intelligence.” It uses text, images, video, and 3D data to perform tasks like planning a robot’s movements or changing the camera position in a photograph.
  • The rumors that NVIDIA would buy Hugging Face are true. NVIDIA is hoping for a proliferation of models that will run on its hardware, and the company promises that Hugging Face will remain a neutral platform, without favoring one model over another.

Software development

Agents are starting to delegate to, and coordinate with, other agents rather than working solo. Claude Code can break a task apart and hand pieces to other Claude Code instances, Muse Code lets sessions message each other, and Google’s AX orchestrator exists purely to wire up sandboxes and control communications for swarms of agents doing a task together. That shift is pushing developers to rethink what a source repository needs to record, and to start asking how much all this delegation costs.

  • Now that Jev has caught everyone’s attention, what can you build with it? Jevmem is a memory manager that hooks into Claude Code and prunes the context at every conversational turn.
  • AX is a new agent orchestrator from Google. It isn’t an agent; it’s intended to coordinate many agents to complete a task. It creates sandboxes, wires up Git repos and other resources, and controls outbound communications.
  • The latest version of Claude Code can manage Claude projects, breaking a task into subcomponents and delegating the subtasks to other Claude Code instances. Another important change is the ability to read AGENTS.md if CLAUDE.md isn’t available.
  • Google’s CC agent is designed for families. Family members can share data with CC, which has its own user account. It could be used for filling out forms, synchronizing calendars, and other common tasks.
  • What will replace GitHub? There’s a growing consensus that we need different kinds of source repositories to deal with the agent-assisted software development. In addition to changes to code, it’s important to record the conversations between the developer and the agent, the architectural decisions, and many other artifacts that don’t make it into traditional source control.
  • Anthropic is merging its Claude Cowork and chat products. Anything users type in a chat session is seen by Cowork, and vice versa. Some sensitive information (health, politics, and gender) is excluded. The feature is on by default but can be disabled in settings, and memory isn’t shared with Claude Code. The company also released Claude Docs and Slides.
  • Claude Money is a new feature that will allow users to connect their bank accounts to Claude for analysis. The product appears to be similar to a product from OpenAI.
  • OpenAI’s Agents API is now in public beta. It allows compaction, session orchestration, and tool use, and it can be deployed in OpenAI’s sandbox, a cloud provider’s sandbox, or the developer’s hardware.
  • Meta has released Muse, its AI agent. Muse is a “personal agent” designed for tasks like shopping, filling in forms, and dealing with customer service. It has its own secure credential store, so data like passwords and credit card numbers are never sent offsite.
  • Some open source projects are shutting down external pull requests, which are largely AI-generated. In some cases, the developer team is using its own agents to create and manage PRs; some are using AI agents to triage external PRs.
  • Meta has launched Muse Code, another competitor to Claude Code. One important new feature is the ability to send messages to other Muse Code sessions, allowing agents to coordinate on complex problems.
  • AI providers appear to be moving toward outcome-based pricing, at least for major corporate customers. Rather than billing by token, customers are billed for completed tasks. That approach begs the question: When is a task completed?
  • Now that organizations are concerned with AI budgets, the question of how to evaluate the cost of different models and agents becomes important. What should platform teams measure?
  • ChatGPT Work was designed to compete with Claude Cowork, Microsoft Copilot Cowork, Muse Code, and other agents designed for noncoders. Simon Willison shows how Work goes beyond its competitors. It can perform tasks on the web for users, even logging in to websites without sending usernames and passwords to OpenAI; it can execute code with full internet access; it can build and deploy a web application. Whether these features are also risks is an open question.
  • Anthropic has given Claude Desktop access to a Chromium-based browser that’s built into Cowork, eliminating the need for a Chrome plugin when Claude needs to browse the web.
  • TimeLord is a short Python program that, given a string up to 1,000 characters long, produces a seed for Python’s pseudo-random number generator so that repeated calls reproduce the text. It’s a surprisingly simple hack, though not a statement about randomness or the quality of Python’s PRNG.

Security

OpenAI’s experiment that attacked Hugging Face is the gift that keeps giving, but the past month has had plenty of news about more conventional attacks, many aided by AI. Security has always been a game of whack-a-mole, in which vulnerabilities are discovered and exploited as fast as defenders can patch them. AI is an important tool for defenders, and it’s constantly improving, but it’s still behind attackers, especially given the limitations placed on frontier models and the unlimited persistence that attacking agents exhibit.

  • OpenAI has postponed the release of GPT-6.1 Astra because it failed its safety tests.
  • NVIDIA has announced its Open Agent Safety Platform. The reference implementation includes NVIDIA OpenShell, which has been enhanced with a policy prover, and NVIDIA Sentry, a service that runs on NVIDIA DPUs.
  • The Felony Bench lists known attacks by agents from the major AI labs against third parties. We don’t know if the Bench will be kept up-to-date, but tens of thousands of security incidents involving OpenAI and Anthropic are now being investigated.
  • Following through on Dario Amodei’s call to control the speed of frontier model development, Anthropic, OpenAI, and Google are creating a standards consortium for governing the process of AI development. Meta, xAI, Microsoft, and the Chinese labs are all notably absent.
  • Agents need their own identity. Unlike the long-term identities we’re used to, agents need a short-lived identity tied to a revocable certificate and that limits access to resources appropriate for the job. That’s not all of agent security, but it’s a table stakes.
  • Anthropic has published a lengthy report on the misuse of its systems by threat actors. Daniel Meissler has published a summary, digesting Anthropic’s report into 117 findings.
  • A malicious NPM malware package works by hiding malicious code in the package itself (indexed-btree) rather than simply attacking the install script. This technique makes it significantly harder to detect.
  • An attack against the RSA algorithm allows forging of signatures in some situations. The attack was invented in 2007; this is the first public implementation.
  • Fake CAPTCHA pages are being used to spread malware. Victims are frequently sent to those pages when they respond to a phish.
  • Hugging Face has volunteered to audit AI labs for safety and alignment with human values.
  • In an experiment designed to test AI alignment, DeepMind found that, out of 100 agents, 14% were willing to cheat, 25% were “whistleblowers” that reported cheating, and the remainder didn’t notice.
  • Threat actors are building frameworks for AI agent-enabled attacks. A human in the loop is no longer needed. Fully automated attackers don’t appear to be using zero-days yet; they’re relying on known vulnerabilities.
  • OpenAI autonomous AI agents were found communicating with each other via publicly accessible Wikis, possibly to collaborate on a benchmark.
  • OpenAI has stated that its unreleased Astra model has reached the “Critical” cybersecurity threshold, which means that it can find new vulnerabilities and run exploits against well-protected systems. Now that Astra is released, access to its cybersecurity capabilities has been limited.

Infrastructure and Operations

Individuals, corporations, and even nations all face a similar problem: keeping their infrastructure under control. At a minimum, control means keeping data on a laptop, corporate server, or data center; at the other end of the spectrum it means eliminating dependencies on software and services from another nation. Any organization working through an AI transformation has to evaluate its entire stack: What do they need to control, and what can they safely delegate to others?

  • DAWO is a community that’s building an open source “workspace” to support digital sovereignty for the Dutch government. The stack will include AI, an operating system based on NixOS, cloud services, and collaboration tools.
  • Cohere now offers a confidential computing platform for artificial intelligence. The company claims that customer data is never visible to Cohere itself or any cloud providers that are in use; data is processed on GPUs whose memory is encrypted and isolated.
  • Perplexity has announced Hybrid Compute, a feature that allows it to run models and use files and tools directly on a user’s Mac. The company claims that sensitive data will never leave the user’s computer.

Hardware

It’s too easy to view consumer devices as innocuous things that sit around and do their job silently. Recent devices include cameras, microphones, and even EEG sensors that are constantly collecting data. Where is that data sent, how is it used, and who might have access to it? These questions need to be asked more often.

  • LG Smart Televisions have been found to record conversations and other audio, even while turned off. The conversations are sent back to LG. If the set is disconnected from the network, it will attempt to find open WiFi access points to deliver its data.
  • In part because of backlash against Meta’s camera-enabled glasses and their abuse, its AI glasses now come with or without a camera, and can be used as hearing aids. Well-documented abuse aside, virtual reality will only succeed if there are fashionable, easily wearable products.
  • Headphones, earbuds, and other devices equipped with EEG sensors are appearing on the market. They’re advertised for monitoring fatigue, monitoring sleep, and similar applications. It’s time to ask what happens at the interface between neurology and AI.
  • Microduck is a small bipedal AI-driven robot. It’s trained in simulation with open source software, and the model that results can be shared on Hugging Face. It’s affordable and is available for preorder now, shipping by Christmas.

Web

  • Cloudflare now supports HTTP Vary, which allows servers to serve different kinds of files at the same URL. This is the “ugliest part” of the HTTP standard. It makes caching very difficult, and it probably should be avoided.
  • WebMCP is a proposed standard that gives websites a small API to register tools that agents can discover and call. It was developed by Google and Microsoft.
  • A new Twitter? Operation Bluebird is relaunching Twitter, the service bought by Elon Musk and renamed X.

Biology

  • Anthropic has built a biology lab for experimenting with AI-enabled drug development. Claude assisted in the discovery of an enzyme that might be able to perform CRISPR-like gene editing.
  • To improve its training data for biological applications, the OpenAI Foundation (OpenAI’s nonprofit parent organization) is buying data from failed biotech companies.
  • Google has released AlphaGenome Atlas, a database of every possible single letter change to human DNA, and what that change will do.


Read the whole story
alvinashcraft
2 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

JetBrains Joins the Open Source Security Foundation

1 Share

JetBrains has joined the Open Source Security Foundation (OpenSSF) as a general member. OpenSSF is a cross-industry initiative of the Linux Foundation that works to secure the software supply chain on which the industry depends. The foundation announced our membership today, together with other new members, at OpenSSF Community Day Europe in Prague, co-located with Open Source Summit Europe.

OpenSSF is home to projects that many teams already rely on, including Sigstore for signing and verifying software and SLSA for supply chain integrity. Its AI/ML Security Working Group publishes guidance on signing machine learning models and on writing safe instructions for AI code assistants.

Open Source Security Foundation General Member

Why we are joining

As coding agents take on more development work, code becomes cheaper to produce but more expensive to verify, and much of that verification is security work. Developers need to know where their code and dependencies come from, and AI adds new questions about how to trust the models and assistants that help write it.

JetBrains Air is how we are working on these questions in our own products, helping developers understand and verify what agents produce. OpenSSF is where we can work on the same questions out in the open, with the wider industry.

“Software development is at an inflection point. AI is changing how software is built and creating new security challenges, making it more important than ever that developers can understand, verify, and trust the software they produce.

JetBrains has supported professional software development for more than two decades, and we believe staying ahead of these challenges is best done collaboratively and in the open. OpenSSF brings together some of the strongest expertise in the industry, and we are glad to join the community and help shape the future of secure software development.”

– Katherine Druckman, Head of Community and Partnership Engagement, JetBrains

We want to stay ahead of the security questions AI raises, and help guide how the industry answers them. OpenSSF is the right place to do both.

Looking ahead

Developers trust JetBrains tools with their code every day. We take that trust seriously, and supporting the community working to make software more secure is part of how we aim to keep earning it. We look forward to working alongside the rest of the OpenSSF community.

Read the whole story
alvinashcraft
2 minutes ago
reply
Pennsylvania, USA
Share this story
Delete

Critter Stack Update with Jeremy Miller

1 Share
The Critter Stack is growing! Carl and Richard talk with Jeremy Miller about the latest Critter Stack updates, including Marten, Wolverine, and more! First up is the stack's expansion with Polecat and Fisher, versions of Marten built for SQL Server 2025 and SQLite, respectively. Jeremy also talks about the evolution of event sourcing and how the framework continues to advance to take advantage of new approaches to managing fast, timely data flows. The conversation also digs into how AI is impacting frameworks, including the continued need for reliable frameworks so you can focus on providing value to your customer. But also in the AI vein is the AI Skills for the Critter Stack - so your AI development tools can work as effectively as possible, and get your application built!



Download audio: https://dts.podtrac.com/redirect.mp3/api.spreaker.com/download/episode/75608463/dotnetrocks_2023_critter_stack_update.mp3
Read the whole story
alvinashcraft
3 minutes ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories