Sr. Content Developer at Microsoft, working remotely in PA, TechBash conference organizer, former Microsoft MVP, Husband, Dad and Geek.
161066 stories
·
33 followers

AI Agent ROI Framework

1 Share

Introduction 

AI agents are rapidly becoming a core component of enterprise transformation strategies. Organizations are deploying agents to assist employees, automate business processes, improve customer experiences, accelerate decision-making, and scale operations. Yet many organizations still face a fundamental challenge.

How do we prove the business value of AI agents? 

While innovation and experimentation are important, long-term success requires a disciplined approach to measuring outcomes, managing costs, and continuously optimizing investments. Organizations that fail to establish clear ROI measurements often struggle to justify continued funding, prioritize future investments, or scale successful solutions. 

An effective AI Agent ROI Framework helps leaders connect technology investments directly to business outcomes. It enables organizations to identify the right use cases, understand cost drivers, quantify value, establish governance, and create a culture of continuous optimization. 

This article presents a practical framework that can be used across industries and business functions to define, measure, and maximize AI agent return on investment. 

Start with the Right Use Cases 

One of the biggest mistakes organizations make is evaluating AI projects solely through a technology lens. Successful organizations begin with business outcomes. 

Before building an AI agent, leaders should evaluate opportunities across six dimensions: 

  • Reduction of operational overhead 
  • Resource optimization 
  • Scalability improvements 
  • Workforce productivity gains 
  • Customer experience enhancement 
  • Revenue growth potential 

The highest-performing AI initiatives typically address repetitive, high-volume, and measurable business processes where automation can quickly generate value. 

Examples include: 

  • Employee support assistants 
  • Knowledge retrieval agents 
  • Customer service agents 
  • Sales enablement assistants 
  • Compliance review agents 
  • IT operations agents 

Organizations should prioritize quick wins that provide: 

  • Low implementation complexity 
  • High business impact 
  • Short time-to-value 
  • Strong executive sponsorship 
  • Clear success metrics 

Establishing early wins creates momentum, increases adoption, and generates measurable evidence that supports broader AI investments. 

Understanding the Cost Structure of AI Agents 

Accurate ROI calculations require a complete understanding of costs. Many organizations underestimate the true investment needed to deploy and operate AI agents successfully. 

AI agent costs generally fall into five categories.

1. Infrastructure and Platform Costs

These costs include: 

  • AI model consumption 
  • Compute resources 
  • Storage 
  • Networking 
  • Hosting platforms 
  • Development environments 

Organizations should select the simplest architecture that satisfies business requirements. Overengineering often becomes one of the largest sources of unnecessary spending.

2. Development and Integration Costs

Building an agent often requires: 

  • Solution design 
  • Business analysis 
  • Prompt engineering 
  • System integration 
  • API development 
  • Security implementation 
  • Testing and evaluation 

These investments form the "cost to achieve" portion of ROI calculations.

3. Data Preparation Costs

Data remains one of the most significant contributors to AI project success. 

Poor data quality increases costs through: 

  • Additional token usage 
  • Lower response quality 
  • Greater compliance risk 
  • Increased maintenance effort 

Organizations should invest early in data cleansing, deduplication, classification, and governance.

4. People and Skills Costs

Successful AI implementations require cross-functional collaboration among: 

  • Business stakeholders 
  • AI engineers 
  • Data engineers 
  • Developers 
  • Security professionals 
  • Governance teams 
  • User experience specialists 

Many organizations underestimate the ongoing talent investment required to operate AI solutions at scale.

5. Ongoing Operational Costs

Agent deployment is not the end of the journey. 

Ongoing expenses include: 

  • Monitoring 
  • Evaluation 
  • Model updates 
  • Prompt optimization 
  • Governance reviews 
  • Support operations 

Many enterprises find that annual operational expenses represent a meaningful percentage of the original implementation investment and should be included in ROI planning.

Defining Business Value 

A common challenge when measuring AI ROI is focusing only on labor savings. 

The most successful organizations evaluate value across multiple dimensions. 

Productivity Impact 

Productivity improvements include: 

  • Faster task completion 
  • Reduced manual work 
  • Improved employee efficiency 
  • Reduced process latency 

Typical measurements: 

  • Hours saved 
  • Tasks automated 
  • Cycle-time reduction 
  • Increased throughput 
Financial Impact 

Financial value may include: 

  • Revenue growth 
  • Cost reduction 
  • Resource optimization 
  • Reduced outsourcing spend 

Examples include higher conversion rates, increased deal velocity, or lower support costs. 

Risk Reduction 

AI agents can provide value through: 

  • Compliance improvements 
  • Better policy adherence 
  • Reduced operational errors 
  • Faster issue detection 

Risk mitigation often generates significant value even when direct cost savings are difficult to observe. 

Strategic Value 

Not all benefits can be measured immediately in dollars. 

Strategic outcomes may include: 

  • Improved customer experiences 
  • Better decision-making 
  • Increased innovation 
  • Brand differentiation 
  • Workforce empowerment 

Organizations should measure both financial and strategic outcomes to capture a complete picture of value.

Example ROI Statement (for a Business Requirement Document) 
  • Manual process: 15 minutes per task 
  • Expected agent time saved: 10 minutes per run 
  • Expected runs per month per user: 10 runs/month per user 
  • Number of Users: 20 users ( whole team) 
  • User Adoption: 15 users for Phase 1 (first month after Agent Onboarding,75% adoption rate). Planned to expand to 100% adoption rate at Phase 2, with 20 users.  
  • Monthly time saved: 
    • At Phase 1: 18,000 min per year (300 hours)  
    • At Phase 2: estimated to 24,000 min saved per year (400 hours) 
  •  Additional benefits: Reduced SLA breaches, improved data accuracy, automated documentation. 

This becomes your measurable target for post-launch monitoring.  

Building an AI Agent ROI Model 

A practical ROI calculation includes three primary components: 

Cost to Achieve 

The investment required to build and deploy the solution. 

Examples: 

  • Development 
  • Licensing 
  • Integration 
  • Infrastructure 
  • Training 
Cost to Maintain 

The ongoing investment required to operate the solution. 

Examples: 

  • Model consumption 
  • Monitoring 
  • Governance 
  • Enhancements 
  • Support 
Benefits Generated 

The measurable value created by the AI agent. 

Examples: 

  • Labor savings 
  • Revenue improvements 
  • Risk avoidance 
  • Productivity gains 

A simplified ROI formula can be expressed as: 

 However, mature organizations move beyond simple ROI and evaluate investments across multiple years. 

 Example: Zava's Customer Service Agent - 3-Year Projection 

  • Cost to Achieve = $80,000 
  • Cost to Maintain = 20,000/year*3years = 60,000 
  • Total Benefits over 3 years = $200,000 

When used with the formula to calculate ROI this achieves a result of 42.86% 

ROI = (60,000 ÷ 140,000) × 100 = 42.86% 

Interpret the result 

  • ROI > 0% → Profitable investment 
  • ROI < 0% → Loss-making investment 

A 42.86% ROI means the agent returns nearly 43 cents for every dollar spent over 3 years. 

Looking Beyond ROI with NPV 

Many AI investments generate value over several years. 

Net Present Value (NPV) provides a more comprehensive evaluation by accounting for the time value of money. 

 

Where: 

  • CFt = cash flow in year t 
  • r = discount rate 
  • n = number of years 
  • I = initial investment 

Let's say: Zava's Return Fraud Detection Agent - 5-Year NPV 

  • Initial Investment = $100,000 
  • Annual Cash Flow Impact (positive) = $30,000 
  • Time Horizon = 5 years 
  • Discount Rate = 8% 

Then: 

 

The 5-year discounted cash flow impact (NPV) is $19,781 dollars. Without applying the discount rate, the cash flow impact would be 50,000 dollars. 

Interpret the Result 

  • NPV of the cash flow impact > 0 → Investment could be approved (subject to budget constraints) 
  • NPV of the cash flow impact < 0 → Investment should be rejected 

This approach is especially useful for enterprise-wide AI platforms, multi-agent ecosystems, and large-scale digital transformation initiatives. 

When finance teams evaluate strategic AI investments, NPV often becomes a more meaningful indicator than ROI alone. 

Understanding sensitivity analysis for AI investments – Managing Uncertainty 

AI investments involve uncertainty.  

Identify key variables 

Start by listing the assumptions that significantly influence ROI. For AI agents, these might include: 

  • Adoption rate is the #1 variable (for example, % of users actively using the AI agent) 
  • Development cost (for example, agents' building and change management) 
  • Operational cost (for example, maintenance, updates, cloud usage) 
  • Performance outcomes (for example, time saved, errors reduced, revenue generated 

Establish a baseline scenario 

Define the expected values for each variable based on current data or forecasts. For example: 

  • Adoption rate: 70% 
  • Development cost: $90,000 
  • Operational cost: $12,000/year 
  • Annual savings: $48,000/year 

Model different scenarios: 

Scenario 

Adoption Rate 

Dev Cost 

Annual Savings 

NPV 

Optimistic 

90% 

$70K 

$60K 

$140,000 

Baseline 

70% 

$90K 

$48K 

$50,000 

Conservative 

50% 

$110K 

$35K 

$1,000 

Worst-case 

30% 

$130K 

$20K 

-$50,000 

Governance Frameworks That Support ROI 

Organizations attempting to scale AI should establish governance structures that balance innovation with operational discipline.

AI Center of Excellence (AI CoE) 

A strategic hub for AI strategy, governance, and standardization 

Building the CoE — 5 Steps 

  1. Secure executive sponsorship — budget, authority, credibility 
  2. Appoint a CoE leader — single point of contact for AI strategy 
  3. Assemble cross-functional team — data scientists, engineers, ethicists, business leaders 
  4. Determine organizational placement — central vs. embedded 
  5. Define operating model — centralized early, advisory as maturity grows 
FinOps for AI 

Traditional cloud cost management principles are increasingly being extended to AI workloads. 

AI FinOps focuses on: 

  • Cost transparency 
  • Consumption accountability 
  • Usage optimization 
  • Budget forecasting 
  • Resource right-sizing 

Organizations should establish visibility into costs at the level of agents, business units, products, and environments. 

GenAI Operations (GenAIOps) 

Continuous evaluation and operational excellence have become critical capabilities. 

Core practices include: 

  • Prompt lifecycle management 
  • Evaluation frameworks 
  • Observability 
  • Performance monitoring 
  • Quality tracking 
  • Automated improvement loops 

Together, AI CoE, FinOps, and GenAIOps provide the governance foundation necessary to maximize long-term ROI. 

Continuous Measurement and Optimization 

Many organizations calculate ROI once and never revisit it. 

High-performing organizations treat AI ROI as an ongoing process. 

Recommended metrics include: 

Business Metrics 

  • Revenue impact 
  • Cost savings 
  • Productivity gains 
  • Risk reduction 

Adoption Metrics 

  • Active users 
  • Utilization rates 
  • User satisfaction 
  • Repeat usage 

Operational Metrics 

  • Response quality 
  • Success rate 
  • Resolution rate 
  • Escalation rate 

Financial Metrics 

  • Cost per interaction 
  • Cost per task 
  • Cost per outcome 
  • Budget variance 

Continuous visibility helps organizations identify underperforming agents, optimize successful ones, and prioritize future investments. 

Recommended AI Agent ROI Dashboard 

A centralized dashboard turns this framework into an operating habit rather than a quarterly exercise. Track four groups of KPIs in one place: 

  • Financial ROI %, cost per interaction, cost per task completed, revenue influenced 
  • Productivity hours saved, tasks automated, throughput increase 
  • User experience adoption %, satisfaction score, retention rate 
  • Governance accuracy rate, escalation rate, compliance score, responsible AI metrics 

A Practical AI Agent ROI Maturity Model 

Organizations typically progress through four stages: 

Level 1: Experimentation 

  • Individual pilots 
  • Limited measurement 
  • Basic cost tracking 

Level 2: Managed 

  • Defined KPIs 
  • ROI modeling 
  • Governance processes 

Level 3: Optimized 

  • Organization-wide standards 
  • Continuous monitoring 
  • FinOps practices 

Level 4: Value-Driven AI 

  • Real-time ROI tracking 
  • Portfolio management 
  • Strategic investment optimization 
  • AI treated as a business asset 

The most successful organizations operate at Levels 3 and 4, where AI investments are continuously evaluated and optimized based on measurable outcomes. 

Conclusion 

The future of enterprise AI will not be determined solely by model performance or technological innovation. It will be determined by an organization's ability to create measurable and sustainable business value. 

A successful AI Agent ROI Framework begins with selecting the right use cases, understanding the true cost structure, defining business outcomes, establishing governance, driving user adoption, and continuously optimizing performance. 

Organizations that embrace ROI-driven AI strategies will be better positioned to justify investments, scale successful initiatives, and maximize the long-term value of their AI agent ecosystem. 

The question is not whether AI agents can create value. The question is whether organizations have the discipline and framework required to measure, manage, and maximize that value at scale. 

 

References

 

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Microsoft ODBC Driver 18.7.1: smaller vectors, easier configuration, and better cloud routing

1 Share

Microsoft ODBC Driver 18.7.1 for SQL Server is now generally available.

This release continues the work we started in 18.6 to support modern application patterns while making the driver easier to deploy and operate. It adds float16 vector support, brings familiar connection-string names to ODBC, improves routing for Azure SQL Database Hyperscale, expands platform coverage, and removes a Windows installation dependency.

It also includes a substantial set of reliability, security, and diagnostic fixes. Many of those changes happen below the application layer.

More compact vector workloads

ODBC Driver 18.6.1 introduced support for the SQL Server vector data type using 32-bit floating-point values. Version 18.7.1 adds support for float16 vectors.

A float16 element uses two bytes instead of the four bytes required by float32. For applications working with large embedding collections, that can reduce the amount of vector data stored and moved between the application and the database. The right precision depends on the model and workload, but applications that can use half-precision vectors now have that choice through ODBC.

This matters because vector support is becoming part of the normal database application stack. You should not need a separate connectivity path just because an application combines relational data, business data, and embeddings. Adding float16 support brings ODBC forward with the vector capabilities being added across SQL Server and Azure SQL.

Connection strings that travel more easily

ODBC has accumulated its own connection-string vocabulary over many years. Some settings use different names in ODBC, JDBC, OLE DB, and SqlClient even when they configure the same behavior.

Version 18.7.1 accepts four additional connection-string keywords:

  • MultipleActiveResultSets
  • FailoverPartner
  • WorkstationID
  • ConnectTimeout

The first three are aliases for existing ODBC settings. For example, MultipleActiveResultSets maps to MARS_Connection, while WorkstationID maps to WSID. Existing connection strings continue to work.

ConnectTimeout maps to the ODBC login timeout. It supports the same behavior as the underlying ODBC setting, including zero for an infinite timeout and a default of 15 seconds.

These additions reduce the small but persistent differences you encounter when moving configuration between Microsoft SQL drivers. Shared configuration systems, deployment templates, and migration tools can use more consistent names instead of maintaining driver-specific translations for common settings.

We also spent some time on details that tend to cause production surprises: case-insensitive matching, precedence when an alias and canonical name both appear, invalid values, timeout limits, and DSN interaction.

Better routing for Hyperscale read workloads

Azure SQL Database Hyperscale named replicas provide independent read scale for applications with large or isolated read workloads. ODBC Driver 18.7.1 adds load-balanced routing for named-replica reader endpoints.

Applications can connect through the reader endpoint and allow the service and driver to handle routing across the available read capacity. That makes the endpoint more useful for workloads such as reporting, analytics, and read-heavy application services without requiring applications to manage individual replica destinations themselves.

The driver already sits at the point where connection intent becomes a physical connection. Supporting this routing behavior there keeps replica topology out of application code.

Less setup on Windows

The Windows package no longer requires the Microsoft Visual C++ Runtime to be installed separately.

That removes a prerequisite from new machines, container images, automated build agents, and managed desktop deployments. It also reduces one of the common differences between a machine where an application was built and a clean machine where it is installed.

The change is small from an application-code perspective. For deployment owners, it means fewer moving parts and one less prerequisite to diagnose.

More Linux distributions

Version 18.7.1 adds support for:

  • Alpine Linux 3.23
  • SUSE Linux Enterprise Server 16
  • Ubuntu 26.04

Platform support is more than producing an RPM, DEB, or APK. The driver must install cleanly, connect successfully, upgrade from prior versions where packages are available, and coexist with ODBC Driver 17 on supported configurations.

The installer automation used for this release covers AMD64 and ARM64 variants across RHEL, Azure Linux, Ubuntu, Debian, Alpine, and supported SUSE environments. Windows coverage includes Windows 11 and Windows Server 2019, 2022, and 2025, with fresh installation, upgrade, coexistence, and MSI repair scenarios where applicable.

Ubuntu 26.04 currently receives fresh-install validation because previous packages are not yet available from the Ubuntu 26.04 Microsoft package repository. The test pipeline records that distinction instead of treating unsupported upgrade combinations as covered.

Reliability work below the application

Database drivers operate on untrusted network input, coordinate asynchronous operations, manage native memory, and translate between platform APIs and the Tabular Data Stream protocol. Small mistakes in those paths can produce failures far away from the code that caused them.

The 18.7.1 release addresses several of those cases:

  • Protocol parsing is more defensive when processing malformed LOGINACK and ENVCHANGE tokens.
  • Memory corruption issues were corrected in the SQL Server Network Interface packet pool and on Linux ARM64 and macOS ARM64.
  • Multiple Active Result Sets connections now clean up memory correctly when a connection ends abruptly.
  • Asynchronous timeout handling was corrected for zero-length partially length-prefixed data and data-classification tokens.
  • OpenSSL errors left on a calling thread are handled correctly.
  • XA distributed transactions recover more reliably from SQL Server connectivity failures.
  • Always Encrypted performs less redundant logging while acquiring Azure Key Vault tokens.
  • Tabular Data Stream packet tracing reports the correct byte count for overlapped named-pipe writes.

Most applications will never encounter the exact failure conditions behind these fixes. That is the goal. A malformed server response should produce a controlled error. A dropped connection should release its memory. A timeout should behave consistently even when it arrives in the middle of an unusual protocol sequence.

Get ODBC Driver 18.7.1

Microsoft ODBC Driver 18.7.1 for SQL Server is available now for Windows, Linux, and macOS.

If your application uses vectors, Azure SQL Database Hyperscale named replicas, shared connection configuration, or newer Linux distributions, 18.7.1 contains changes you can use immediately. For other applications, the deployment, protocol, memory-management, and diagnostic fixes provide a strong reason to include this release in your normal driver update cycle.

Read the whole story
alvinashcraft
just a second ago
reply
Pennsylvania, USA
Share this story
Delete

Create multimodal applications with OpenAI models in Microsoft Foundry

1 Share

Imagine you are building a home design application. A customer says, “Show me a modern kitchen with dark cabinets and a large center island.” Images appear in seconds. Before the application finishes describing the design, the customer interrupts: “Make it brighter and add natural wood accents.” The application listens, adapts, updates the visuals, and continues the conversation without losing context. OpenAI’s GPT-Live-1 and GPT-image-2.5 models make this type of conversational and creative experience possible, and they are now available in Microsoft Foundry.

Many applications today operate through a sequence of prompts and responses, one after the other, no overlap. The user asks a question, the AI answers, and the interaction starts again. Human collaboration does not work that way. People interrupt, clarify what they mean, react to new ideas, change direction, and build on concepts together in real time.

GPT-Live-1 brings that natural rhythm to AI applications through real-time, full-duplex conversation. The model can listen and speak simultaneously, helping developers create experiences that respond to interruptions, pauses, overlap, and other conversational signals. GPT-image-2.5 extends that interaction into visual creation, allowing applications to generate and edit images as ideas emerge.

From AI responses to AI collaboration

Until now, developers often had to assemble multimodal experiences from separate systems. A voice interface captured a request, another system processed it, an image model generated the output, and the user moved between different screens or workflows to refine the result.

Voice experiences were also largely built around a familiar turn-based pattern: the user spoke, the system waited for silence, processed the request, and then delivered a response. Even when the answer was useful, the interaction could still feel rigid.

GPT-Live-1 changes that dynamic. Instead of treating conversation as a series of isolated requests, the model can continuously evaluate whether to listen, speak, pause, acknowledge the user, respond to an interruption, or call a tool. This makes conversation more than an input method. It becomes the interface through which the user explores an idea and directs an application.

When that real-time conversation is combined with GPT-image-2.5, the application can turn ideas into visual output as the discussion unfolds. The user does not need to translate every thought into a carefully structured prompt or restart the workflow whenever a requirement changes. They can speak naturally, see the result, react, and continue creating.

What is GPT-Live-1?

Voice is becoming an important interface for customer engagement, employee productivity, accessibility, education, and hands-free work. However, many voice applications still depend on a rigid exchange in which one person speaks and the system waits for silence before responding.

GPT-Live-1 is designed for continuous, real-time interaction. Its bidirectional, full-duplex design allows it to process incoming audio while generating speech, creating experiences that can listen and speak at the same time. The model can respond to interruptions, adapt its pacing, recognize pauses, and use conversational timing to determine when to listen or respond.

The result is not simply faster voice AI. It is a different interaction model. Instead of requiring users to adapt their behavior to the application, developers can build applications that adapt to the natural flow of a conversation.

Key capabilities

  • Simultaneous audio input and output: Support continuous conversation without requiring strict turn-taking.
  • Natural interruption handling: Allow users to clarify, redirect, or add information while the application is speaking.
  • Conversational timing: Respond to pauses, pacing, and other signals that help conversations feel less mechanical.
  • Backchannel responses: Signal that the application is listening without unnecessarily taking over the conversation.
  • Contextual memory: Maintain continuity as users revisit ideas and refine what they need.
  • Multimodal inputs: Work with voice, text, and images to support richer application experiences.

How developers are using GPT-live-1

GPT-Live can support voice-first experiences wherever natural timing, continuous listening, and fast responses matter. Developers can combine its real-time speech capabilities with tools and application context to create experiences that move beyond rigid turn-by-turn interactions.

  • Customer service and contact centers: Build voice agents that answer questions, guide customers through tasks, respond to interruptions, and call tools to retrieve account or order information.
  • Voice-enabled productivity assistants: Let employees capture notes, find information, update records, or complete routine workflows through hands-free conversation.
  • Conversational commerce and discovery: Help customers refine product, travel, or service searches by describing preferences naturally and adjusting them throughout the conversation.
  • Education and coaching: Create interactive tutors, language-practice partners, and training simulations that adapt explanations, pacing, and follow-up questions in real time.
  • Accessibility experiences: Add responsive voice interaction to applications for people who prefer or rely on speech as an alternative to typing and screen-based navigation.
  • Field and frontline support: Deliver hands-free guidance for technicians, warehouse teams, and other mobile workers who need immediate answers while completing physical tasks.
  • Multilingual experiences: Support conversational applications that serve users across languages and use cases such as travel assistance, public services, and global customer engagement.

Audio models in action

CoStar Group is a leader in demonstrating how conversational audio can reshape a familiar digital experience. On Homes.com and Apartments.com, the company has developed an AI-powered housing-search experience that enables buyers and renters to describe what they want conversationally rather than relying only on fixed filters.

"We saw real-time conversational AI as the future of how people search for homes, and we wanted to lead that shift rather than follow it. Early adoption of Azure OpenAI’s Realtime models through Microsoft Foundry gave CoStar Group a foundation to bring that experience to millions of homes.com users.” - Andy Ventura, Vice President Applied AI and Enterprise Architecture

Using Azure Open AI models in Microsoft Foundry, CoStar designed the solution for responsive audio interaction, strong security, and the ability to scale across a marketplace used by more than 100 million visitors each month. A buyer can progressively refine a search through dialogue and explore details that can be difficult to express through traditional search fields, from neighborhood preferences to property and Matterport-powered visual insights.

Learn more in the CoStar Group customer story

What are GPT‑Image‑2.5 Flare and GPT‑Image‑2.5 Sunburst?

The GPT‑Image‑2.5 family gives developers two options for building image generation and editing into user-facing applications. GPT‑Image‑2.5 Flare is a fast, high-throughput model designed for most production workloads. As the smaller model, Flare delivers improved image quality, editing accuracy, and responsiveness for content creation, marketing assets, product experiences, visual search, and image generation at scale. GPT‑Image‑2.5 Sunburst is a premium model optimized for creative workflows that require greater precision and control. Sunburst is designed for high-fidelity campaign assets, product imagery, and other polished visual experiences where quality takes priority over generation speed. Together they bring:

  • Significantly faster generation and editing. GPT‑Image‑2.5 Flare delivers higher-quality images than GPT‑Image‑2 at 50% lower latency, helping keep users in a live, conversational iteration loop.
  • More accurate editing. Update targeted elements while maintaining the rest of the image.
  • Stronger multi-turn editing. Refine images through successive instructions without losing consistency or image quality.
  • Better instruction following. Interpret complex visual instructions, layouts, and stylistic directions more accurately.
  • Two models for different workflows. Use Sunburst for premium creative and editing workflows and Flare, the smaller model, for faster iteration.

See it in action:

Below is a series of images that demonstrates the improvements in multi-turn editing with GPT-images 2.5 Sunburst. The sequence shows the model responding to a series of editing instructions over multiple turns. As changes accumulate, the model maintains visual consistency, preserves details from earlier edits, and incorporates new requests without degrading image quality. These examples highlight how GPT-images 2.5 Sunburst enables more reliable iterative editing workflows, making it easier to refine assets through a sequence of targeted updates.

Prompt: A pair of running shoes placed on a forest trail at sunrise, sitting on packed dirt scattered with fallen leaves and pine needles. Soft golden sunlight filters through tall trees in the background, casting long shadows across the path. Morning mist hovers low between the trunks. Shallow depth of field, the shoes in sharp focus, photorealistic, natural color palette, cinematic lighting.

Base image

Edit 1 - lighting

Edit 2 - color

Edit 3 - composition

 

For developers and creators, this means less time recreating assets and more time refining them, with the confidence that each edit will build on the last.

How developers can use GPT‑Image‑2.5 models

Retail and media companies can use GPT‑Image‑2.5 Flare to accelerate content creation and production across high-volume digital channels, while GPT‑Image‑2.5 Sunburst supports premium creative and editing workflows that require tighter control across successive edits. Commercial teams can move from an initial concept to channel-ready variations in an interactive workflow without treating every adaptation as a separate production cycle.

  • Scale retail campaign creative. A retailer can create and refine seasonal product imagery for websites, retail media, email, social, and stores, testing settings and promotions while keeping the product and campaign direction consistent.
  • Personalize streaming artwork by viewer segment. A streaming service can tailor thumbnails for the same show to viewers’ preferred genres, then test which versions drive more visits and plays without separate photo shoots.
  • Create personalized virtual try-on experiences. An e-commerce retailer can let shoppers upload a photo or select a model to preview clothing and shoes, then adjust color, style, or setting to compare options before buying.

Bring conversation and creation together

Real-time conversation and image generation enable developers to build applications where users can describe ideas, refine them through dialogue, and see results as the conversation evolves. A traveler might explore destinations by describing preferences and reacting to generated imagery. A homeowner could talk through renovation ideas while new designs are generated in real time. A shopper might explain a preferred style and receive visual recommendations that adapt as requirements become more specific. In these experiences, conversation becomes the interface. GPT-Live-1 handles the interaction while GPT-image-2.5 generates and refines visual content, allowing applications to listen, respond, and create within a single workflow.

  • Travel planning: Users can explore locations, accommodations, activities, or itineraries through an ongoing dialogue while simultaneously viewing personalized imagery.
  • Education and learning: Students can discuss ideas with an AI tutor while receiving generated diagrams, illustrations, visual explanations, and supporting content that evolves throughout the lesson.
  • Customer support: Users can verbally describe an issue, provide photos, receive generated visual guidance, and continue asking questions within a single multimodal interaction.

Pricing

Pricing details will be available on the Microsoft Foundry pricing page within the coming weeks.

Model

Modality

Input

Cached input

Output

gpt-image-2.5-sunburst

Image

$8.00

$2.00

$30.00

Text

$5.00

$1.25

--

gpt-image-2.5-flare

Image

$8.00

$2.00

$30.00

Text

$5.00

$1.25

--

GPT-live-1

Audio

$3.00 per hour----

Pricing is per 1 million tokens unless otherwise specified.

Get Started

GPT-Live-1 and GPT-image-2.5 are available today in Microsoft Foundry, giving developers everything they need to build applications that can listen, speak, generate, and edit content within a single multimodal experience.

To stay current on new model releases and discover implementation examples across model families, visit the Microsoft Foundry Model Releases repository , where you'll find release resources, samples, and technical content accompanying new model announcements. Developers looking to deepen their understanding of models and AI application design can also explore the Model Mastery repository, which provides hands-on workshops and learning paths covering partner models, model routing, optimization, and Microsoft Foundry development experiences. These resources are a great way to continue exploring new capabilities across the rapidly expanding Microsoft Foundry ecosystem.

You can also join the community through Model Mondays and the Foundry Discord to learn from product teams, watch technical deep dives, and connect with other developers building on Microsoft Foundry.

Read the whole story
alvinashcraft
16 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

How to choose a FHIR Server: a practical framework for technical teams

1 Share

Every FHIR (Fast Healthcare Interoperability Resources) server can show a green checkmark on a conformance test. That’s about where the similarities end. The honest answer to “which server should we use” depends on your implementation guides, data volumes, terminology needs, and how much of the stack you want to own.

If you’re following our Interoperability campaign, you know the pressure: EHDS (European Health Data Space) and CMS-0057-F, the US interoperability and prior authorization rule, are both pushing organizations to stand up FHIR infrastructure on a fixed timeline, often for the first time. Get the server choice wrong and you’re not fixing a config file, you’re re-platforming mid-deadline. We hear constantly from implementers that this decision gets made faster than it should. A team picks whatever server ships with their cloud contract, or whatever shows up first in a search, and only discovers the gaps once they are three implementation guides deep.

This post walks through the criteria that matter, in the order we’d actually check them, and ends with a scorecard you can run against your own shortlist, whichever vendors are on it.

FHIR is a specification, not a product, and the gap between “conforms to FHIR” and “does what your project needs” is bigger than most teams expect.

A few things surprise people:

  • No server implements the entire FHIR spec. Coverage varies, and features you assume are standard, like specific search modifiers or bulk export, may need to be checked directly against a server’s own capability statement rather than assumed.
  • Support for the newest FHIR version is inconsistent. R4 (2019) is still what most regulations and Da Vinci implementation guides reference, and remains the safest default for production in 2026. R5 support is improving but uneven across vendors.
  • A lower price doesn’t mean lower total cost. The cheapest option on paper is often the one that shifts the most work onto your own team, in the form of custom code, tuning, and troubleshooting you did not budget for. That cost shows up later, not on the first invoice.

None of this means FHIR servers are immature. It means evaluation has to happen at the level of your actual implementation guides and data, not a vendor’s marketing page.

We think about server evaluation across six areas. None are optional, but which matters most depends on your use case.

Deployment model. You have a few options here: Managed cloud, commercial enterprise, or open source and self-hosted. Facade servers, sitting in front of an existing data store like an EHR database, are a fourth pattern worth knowing.

Fit with your current stack. The language, database, and cloud a server is built on decides how much of your team’s existing skills you can reuse.

Conformance and profiling depth. How strictly does the server validate against your implementation guide? Matters for guide-heavy work like Da Vinci PAS, less for simple CRUD.

Data ingestion. How easily can you get data into the server? A REST API is table stakes; check for Subscriptions (FHIR’s pub/sub mechanism) and bulk or legacy ingestion too, like Firely Server Ingest.

Extensibility. Real projects need custom logic. Identity matching, validation, integration hooks. Some servers offer clean extension points; others push it onto a layer you build yourself.

Operational ownership and support. Who tunes performance when it degrades, and answers when something breaks at 2 a.m.? Vendor track record is a proxy for future support.

Cost sits underneath all six, but these questions determine whether the number on the quote is the number you actually pay.

Named vendor lists age fast, and they’re never complete: there’s always another solid option that didn’t make the cut, and someone will always point that out. Use the six criteria above as a scorecard instead. Copy it for every FHIR server on your shortlist, score each one honestly, and let the pattern tell you where the real trade-offs sit for your project.

Criterion Score (1–5) What to check 
Deployment model   Open source and self-hosted, managed cloud, or commercial enterprise? 
Fit with your current stack   Which language, database, and cloud does it expect you to run? 
Conformance and profiling depth   How strictly does it validate against your actual implementation guide? 
Data ingestion   REST API, Subscriptions, bulk import: which of these does it support well? 
Extensibility   Clean extension points, or a layer you have to build yourself? 
Operational ownership and support   Who tunes it when it degrades, and how fast do they answer? 

A word on published benchmarks: request-per-second numbers and price points move quickly and depend heavily on workload shape. The only benchmark that should actually decide your shortlist is one you run yourself, against your own implementation guides and your own data volumes.

Treat the server choice as infrastructure, not a checkbox. It shapes your integration timeline, governance model, and total cost of ownership for years. Expensive to unwind once you’re live.

Test before you commit. Load your actual profiles, ValueSets, and a realistic data volume, and see what breaks. The gap between a vendor’s demo and your production reality is where most surprises live.

If you don’t have a finalized implementation guide yet, which is common for teams preparing for EHDS since the detailed spec isn’t expected until around March 2027, test against the closest reference IG instead, such as US Core, and favor a deployment model you can adjust once the real spec lands.

Involve compliance and business stakeholders early, not just engineering. If EHDS, CMS-0057-F, or a similar mandate is why you’re doing this now, the deadline and audit requirements belong in the evaluation criteria too.

We built Firely Server around this reality. Because our team has been contributing to the FHIR specification since its early drafts and maintains Simplifier.net, we have leaned hard into profiling, validation, and implementation-guide tooling specifically, rather than trying to be the broadest possible platform. That makes it a strong fit for teams whose main challenge is conformance and IG authoring. If your main bottleneck is something else, like raw ingest throughput at massive scale, weigh that against the other criteria above. Knowing which challenge you actually have is most of the evaluation.

The right FHIR server is the one that matches your implementation guides, your data volumes, and how much operational ownership your team actually wants to take on, not the one with the longest feature list or the most confident sales deck. Every scorecard, including the one above, is a starting point for a shortlist. The real answer only shows up once you test candidates against your own project.

It’s not about how many features are on the list either. What matters is whether the ones you actually need work well: fast to implement, reliably performant, with zero unplanned downtime, and backed by support that answers in hours, not weeks, from someone who understands FHIR, not a generic chatbot.

If you are evaluating infrastructure ahead of an EHDS or CMS-0057-F deadline, that testing is worth starting now rather than waiting for every technical detail to be finalized. Talk to us about your implementation or use case, and if you are still getting comfortable with the terminology that comes up in this kind of evaluation, our FHIR glossary is a good place to look things up as you go.

The post How to choose a FHIR Server: a practical framework for technical teams appeared first on Firely.

Read the whole story
alvinashcraft
37 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

Windows Package Manager 1.30.140-preview

1 Share

This is a servicing release of Windows Package Manager v1.30. If you find any bugs or problems, please help us out by filing an issue.

New in v1.30

--ignore-unavailable flag for install

Added a new --ignore-unavailable flag to the install command. When installing multiple packages, this flag allows the operation to continue with the remaining packages instead of failing entirely when one or more packages are not found in the configured sources. This brings the same behavior previously available with import --ignore-unavailable to direct multi-package installs.

Bug Fixes

  • Fixed an issue where winget search --id <msstoreId> could fail to return a Microsoft Store package unless --exact was also provided.
  • Updated NUnit to v4
  • Fixed a crash (0x8000ffff) when using --disable-interactivity with the Resume experimental feature enabled during install operations.
  • Fixed relative path handling for rooted paths.

What's Changed

Full Changelog: v1.30.100-preview...v1.30.140-preview

Read the whole story
alvinashcraft
47 seconds ago
reply
Pennsylvania, USA
Share this story
Delete

v2026.9.4

1 Share

OpenClaw 2026.9.4

Read the whole story
alvinashcraft
52 seconds ago
reply
Pennsylvania, USA
Share this story
Delete
Next Page of Stories