Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
160898 stories
·
33 followers

Beyond the benchmark: How an adaptive approach drives scientific discovery

1 Share

For research and development (R&D) organizations, the promise of agentic AI is not a better one-time answer. It is a new way to explore complex scientific and engineering problems: pursuing multiple hypotheses, validating them against evidence, learning from what does not work, and adapting their approach as new information becomes available.

This unique nature of the agentic discovery process has been a core area of research for Microsoft, and a design principle for Microsoft Discovery, our platform for organizations embracing Frontier R&D.

Measuring adaptive AI for scientific discovery

A new benchmark result shows how that opportunity is becoming real. On Agent’s Last Exam, a demanding evaluation of long-running, tool-using professional tasks, Microsoft Discovery Engine with CLIO (Cognitive Loop via In-Situ Optimization) achieved higher scores than the other agentic harnesses evaluated across three scientific domains: 61.6% in health and medicine, 75.2% in physical sciences, and 64.6% in life sciences.

This result builds on Microsoft’s core research into what makes agentic discovery distinctive. CLIO enables independent reasoning paths to explore a problem, compare and share learning, and resolve the strongest trajectory into a single evidence-backed result. The system can determine when to keep exploring, change strategy, use a different model, or bring a domain expert into the loop.

The CLIO benchmark blog post describes this adaptive reasoning approach in depth. More broadly, this core innovation for scientific discovery, powered by agentic AI, is available to R&D organizations in every industry and the scientific community with Microsoft Discovery not only as a research breakthrough, but as a foundation for real R&D work.

Why scientific discovery requires adaptive reasoning

Many of the hardest scientific and engineering challenges do not have a clearly defined workflow or a known answer. A researcher may need to navigate incomplete evidence, competing objectives, specialized tools, and changing constraints. A materials team may be balancing performance, safety, cost, and manufacturability. A life sciences team may need to connect literature, proprietary data, models, and experimental evidence before deciding what to validate next. An engineering team may need to search a vast design space without sacrificing physical fidelity or traceability.

In these settings, a single model response is not enough. Practitioners need systems that can reason over time, preserve evidence, challenge assumptions, and work within the tools, data, governance, and review processes their experts already use. Just as importantly, they need to understand how a conclusion was reached and where human judgment should enter the process.

Microsoft Discovery was designed as an enterprise platform for agentic R&D, combining the scientific mindset of hypothesis, experimentation, and refinement with the engineering rigor of problem decomposition, structured execution, and reproducibility. CLIO strengthens that foundation with a more adaptive reasoning loop and a diverse model ecosystem, while allowing researchers to use a diverse model ecosystem and multiple reasoning paths.

From benchmarks to real-world impact

The greater opportunity extends beyond benchmark rankings into real research environments. Discovery Engine with CLIO has already supported work that discovered a novel organic redox flow battery. The same approach has potential across design simulation (like for silicon chips), formulation and process optimization (for example in manufacturing and CPG), materials and molecular discovery (which can drive sustainability and drug discovery), and lab automation, areas where organizations need to shorten research cycles, without sacrificing rigor or traceability.

Agentic discovery does not replace scientists and engineers. It expands what they can explore, helps them learn faster from evidence, and gives them a more systematic and transparent way to move from an idea toward an outcome that experts can evaluate and validate.

Realizing the enormous opportunity to redefine R&D requires a platform built for the tools, data, governance, and review processes researchers already use. Microsoft Discovery was designed with that need in mind: to bring agentic discovery to researchers and scientists in R&D organizations across every industry and throughout the scientific community.

We are still early in this journey, but this benchmark milestone demonstrates what becomes possible when AI is built for the way discovery actually happens: iteratively, collaboratively, and adaptively. I look forward to seeing what organizations, researchers, and partners discover next.

Adaptive AI for scientific discovery

Learn how Microsoft Discovery uses adaptive, agentic approaches to explore complex scientific and engineering challenges, helping R&D teams accelerate innovation and uncover new possibilities.

person looking ta the laptop screen in scientific setting

The post Beyond the benchmark: How an adaptive approach drives scientific discovery appeared first on Microsoft Azure Blog.

Read the whole story
alvinashcraft
2 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Java’s age is its AI superpower

1 Share
Ryan welcomes Markus Eisele to the program to talk about why your coding agent should be writing Java.
Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Security Update for SQL Server 2025 RTM CU8

1 Share

The Security Update for SQL Server 2025 RTM CU8 is now available for download at the Microsoft Download Center and Microsoft Update Catalog sites. This package cumulatively includes all previous security fixes for SQL Server 2025 RTM CUs, plus it includes the new security fixes detailed in the KB Article.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Security Update for SQL Server 2025 RTM

1 Share

The Security Update for SQL Server 2025 RTM GDR is now available for download at the Microsoft Download Center and Microsoft Update Catalog sites. This package cumulatively includes all previous security fixes for SQL Server 2025 RTM, plus it includes the new security fixes detailed in the KB Article.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

Security Update for SQL Server 2017 RTM

1 Share

The Security Update for SQL Server 2017 RTM GDR is now available for download at the Microsoft Download Center and Microsoft Update Catalog sites. This package cumulatively includes all previous security fixes for SQL Server 2017 RTM, plus it includes the new security fixes detailed in the KB Article.

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete

How Azure uses AI to turn feedback into improved customer experiences

1 Share

The Challenge: Synthesizing fragmented feedback signals, to improve Azure's experience quality at scale 

Customers experience products and services end to end, but product experiences are often structured around individual service. One team may only know its top issues, while another may see only its own slice of experience.  That structure makes it difficult to identify cross-cutting friction across the broader product experience. 

The feedback signals themselves are also fragmented. Customers share feedback through in-product surveys, support cases, field conversations, and social channels. Most product teams can see only part of that picture, making it hard to distinguish isolated comments from meaningful trends, understand which issues were having the greatest impact, and avoid missing critical feedback. 

For the Azure team, the challenges of fragmentation were amplified by scale. The team processed roughly 10,000 to 15,000 customer feedback reports each month, and synthesizing that feedback required 80 to 100 hours of expert analysis.  Additional effort was needed to translate findings into consistent engineering work items. As feedback volume grew, manual analysis became increasingly unsustainable creating delays in identifying and addressing customer priorities. 

Compounding the challenge was the absence of an effective feedback loop to measure the impact of quality improvements. Teams struggled to justify investments in quality over new features because the return on those investments was difficult to quantify. The absence of a closed-loop measurement system made it difficult to consistently assess the customer impact of quality improvements.  

The team needed a system that could operate across organizational boundaries and across the product development lifecycle: 

  • Identify the most critical customer issues across fragmented feedback channels and product areas. 
  • Convert those insights into actionable engineering work, help teams address issues effectively, and measure outcomes to close the loop. 

The solution needed to preserve team-specific context, maintain auditability, continuously improve through feedback, and keep human experts in control of decisions that require judgment. To address these challenges, Microsoft launched the Great Experiences Matter (GEM) initiative. GEM is designed to analyze feedback signals in aggregate, with access controls and privacy safeguards designed to limit exposure of customer-identifiable information while helping teams identify patterns across channels. 

The Solution: an agentic feedback-to-fix loop with humans in control 

GEM created an AI-enabled feedback-to-fix workflow that connects customer listening, engineering action, and impact measurement across the product development lifecycle. The workflow uses Microsoft Foundry, Azure Data Explorer, Microsoft Fabric, Azure DevOps, and a set of custom agents to transform large volumes of qualitative feedback into prioritized insights and actionable engineering work. 

 

Fig 1. GEM AI-enabled automation workflow 

 

The system closes the loop through two connected motions. 

Find. Agentic workflows remove noise and duplicate reports, assess relevance and actionability, classify feedback against known issues from UX research, cluster related issues, and surface likely root causes. GEM builds on years of deep end-to-end UX research that has identified systemic friction across customer journeys and product boundaries. By continuously triangulating GEM signals with ongoing research, we combine broad, scalable listening with deep human insight to inform a more cohesive Azure experience. 

The results are surfaced through global scorecards for cross-cutting Azure issues and vertical scorecards tailored to individual product teams. 

 

Fig 2. GEM Global scorecard of top issues with Azure, data has been fictionalized to protect intellectual property 

 

Fix. The workflow creates Azure DevOps work items with the customer context, likely reproduction steps, recommended next actions, and an auditable trace of the supporting analysis. To date, 42% of the identified issues have been addressed through engineering action. The team is also extending an AI-assisted engineering workflow, using GitHub Copilot cloud agent, that can generate proposed fixes for straightforward issues. Engineers remain responsible for reviewing, refining, approving, and shipping any changes. 

The architecture is designed for inspection rather than blind automation. The recommendations include an auditable evidence trail allowing reviewers to inspect the source feedback, classifications, supporting references, confidence indicators, and recommendation actions. 

Human expertise enters the system at several points. Researchers shape the issue taxonomies and qualitative grounding. Product teams define ownership boundaries, business priorities, domain-specific vocabulary, and trusted sources that guide agent analysis. Engineers review and act on resulting work items, while leaders use scorecards to inform investment decisions. 

This context is captured in configuration files that evolve alongside the products they support. Teams can add new issue categories, refine keywords, clarify ownership boundaries, or identify trusted research sources. The next analysis cycle automatically incorporates  the updated context without requiring changes to the underlying agents. 

That design creates a "feedback loop for the feedback loop". Teams review the root-cause analyses and work-item quality, identify gaps, and refine their configurations. This enables teams to continuously embed domain expertise into the workflow, improving how feedback is interpreted and prioritized without requiring changes to the underlying infrastructure. Their input improves subsequent runs, helping the system become more precise while preserving local product knowledge. 
 
The agent also maintains access to reports from previous runs and uses tools such as Web IQ and MCP servers to assess whether previously identified issues are improving, still require attention, or can be confidently closed. 

 

 

Fig 3. Example of the vertical feedback agent reasoning through customer feedback to find new work items 

One demonstrated example surfaced customer reports that a networking tool lacked IPv6 validation and support. The workflow generated an engineering work item describing the issue, customer impact, likely reproduction path, and recommended actions. The networking team reproduced the issue, validated the finding, and added it to its backlog. 

The goal is not to remove people from the process. Agents assume much of the cognitive load associated with sorting, clustering, tracing, and drafting, allowing experts to focus on judgment, prioritization, and implementation. Teams remain accountable for what is fixed, what is funded, and what is allowed to ship. 

The Impact: measurable experience gains at Microsoft production scale 

GEM began with manual interventions and is now scaling through AI-enabled workflows. The combined approach has produced measurable results across the Azure Portal and individual product experiences: 

  • The workflow aggregates and analyzes approximately 10,000 to 15,000 feedback reports each month across in-product, support, and social channels. 
  • Automated analysis reduced manual synthesis time by over 95 percent, turning a process that required 110 to 160 hours each month into a workflow that runs in under 60 minutes. 
  • Between October 2025 and April 2026, Azure Portal feedback rates declined by 30%. During that same period, GEM helped teams identify and prioritize experience improvements, creating a clearer link between customer feedback, engineering action, and outcome measurement. 
  • Service-level outcomes also demonstrate how better signals can drive business impact. For example, improvements to VM Connect experiences reduced overall Core Compute support volume by 1-2% per month, resulting in proportionate cost savings. The Azure Growth team increased subscription conversion by 9.1 percent after prioritizing issues highlighted through GEM. 
  • The value goes beyond speed. Leaders gain a more consistent basis for prioritization, and engineering teams receive work that is already connected to customer evidence and impact signals. 

Most importantly, every completed cycle creates new learning. Teams can measure changes in customer feedback and support volumes following improvements, incorporate partner input into future analyses, and continuously refine both the system and the products it helps improve.  

Key learnings and transferable practices 

The GEM experience offers several lessons for teams building agentic systems around complex, qualitative business processes: 

  1. Start with real problems, not AI - Value comes from understanding the genuine business needs and applying AI where it is demonstrably better than existing approaches.  Applying AI without a clearly defined problem often adds complexity without delivering meaningful value. 
  2. Design for the end-to-end workflow - Value comes from connecting insights to the broader business process, including grounding in prior knowledge, prioritization, engineering action, post-fix measurement and reporting. Standalone AI output creates limited value, while an integrated workflow drives outcomes. 
  3. Design for human judgment and accountability - Agents can reduce toil and cognitive load, but researchers, product managers, engineers, and leaders remain responsible for validating insights and determining appropriate actions. 
  4. Ground agents in the knowledge of the teams they serve - Shared models require local context. Editable configuration files allow teams to define ownership, business priorities, releases, examples, and trusted sources without modifying underlying agents. 
  5. Build observability and feedback mechanisms into the agentic system itself - Making analysis inspectable through reasoning traces, source context, and recommendations enables experts to identify gaps, improve outputs, and build trust over time. Build feedback mechanisms directly into the flow of work, making it effortless for users to provide input on the system. 
  6. Tailor outputs to the people making decisions - Executives need trends and investment signals. Researchers need evidence and themes. Engineers need reproducible, actionable work. Effective systems deliver the right information to the right audience. 
  7. Start small and iterate quickly - The AI landscape continues to evolve rapidly. Begin with a well-defined problem, measure outcomes, learn from feedback, and iterate as capabilities mature. 

Looking forward 

GEM continues to scale across the Azure Portal ecosystem. In addition to the global scorecard, vertical scorecards are now live with seven teams, expanding to the top 20 portal extensions representing more than 80% of portal traffic and feedback, with longer-term plans to extend coverage across the entire ecosystem. 

The roadmap includes expanded feedback ingestion, streamlined work-item tracking, AI-assisted remediation workflows, stronger evaluation, and a self-improving architecture. Proposed fixes would remain subject to engineer review, approval, and standard release controls before deployment. GEM is also developing AI-assisted pre-release governance workflows for production code that help identify potential quality issues during development. We will share more about these pre-release workflows in a future post. 

New tools and models will continue to evolve, but the enduring principle remains the same: combine enterprise-scale automation with clear ownership, trusted grounding, and human control.  

For Microsoft, Customer Zero means deploying these systems in real production environments, learning from the complexities, and sharing those lessons broadly. GEM shows what becomes possible when AI does more than summarize feedback. It helps an organization listen, act, measure outcomes, and continuously learn at customer scale. 

 

Microsoft's Customer Zero blog series gives an insider view of how Microsoft builds and operates Microsoft using our trusted, enterprise-grade agentic platform. Learn best practices from our engineering teams through real-world lessons, architectural patterns, and operational strategies for building, operating, and scaling AI-powered systems across the organization. 

Read the whole story
alvinashcraft
3 hours ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories