Not every step in an AI application needs a large language model to generate a response. Many high-volume application steps are decisions: classify an input, select an option, assign a score, verify a condition, prioritize an item, or control a workflow. To suit these tasks, a new class of models have been developed to address these fast, structured, requests.
We are excited to be adding Microsoft-Decision-1 to the Foundry Models Catalog. Our newest, decision model for applications that need to choose among predefined options rather than generate open-ended text.
What are decision models?
As AI applications mature, the industry is moving beyond the idea that every model interaction needs to result in generated content. Developers are increasingly building systems composed of different models, tools, and application logic, each suited to a different part of the workload. Within those systems, many of the most frequent operations are not generation problems at all. They are decisions about what should happen next.
Decision models were designed for this emerging layer of the AI stack. Rather than producing an open-ended response, they evaluate a defined set of possibilities and return structured signals that an application can use to take action. That makes them particularly relevant as developers move toward more dynamic AI applications, where models need to work together and applications must continuously determine where requests should go, which actions should be taken, and when additional processing is needed.
For developers, this creates an opportunity to be more intentional about where and how different models are used. Instead of relying on a general-purpose generative model for every step, developers can match the model to the task, using generative and reasoning models where their capabilities are needed and decision models for focused, repeatable choices. The result is a more modular approach to AI application design, where generation, reasoning, and decision-making become complementary building blocks for creating intelligent systems.
What is Microsoft-Decision-1?
Now available in public preview, Microsoft-Decision-1 is a decision model built on Qwen3.5-9B and designed for applications that need to make fast, structured choices among predefined options. It brings a purpose-built decision capability to the Foundry Models catalog, giving developers another building block for designing AI applications where different steps may benefit from different types of models.
- Built for predefined decisions. Microsoft-Decision-1 supports use cases such as agent controls, model routing, intent analysis, data labeling, AI judging, incident response routing, data validation, and content classification.
- Structured for application logic. The model produces predefined answers and confidence signals that applications can use to select the next action or route uncertain results for further review.
- Designed for lower-latency, lower-cost workflows. Microsoft-Decision-1 is intended for simple, repeatable, and time-sensitive decisions that do not require the generative capabilities or extended reasoning of a larger language model.
You can explore these capabilities in the Foundry Playground. This example below shows Microsoft-Decision-1 checking urgency, selecting a support team, and scoring priority for a customer request.
Microsoft-Decision-1 represents an area of continued model innovation at Microsoft. The team will rebase future iterations of the model on OpenAI models and Microsoft’s MAI models, bringing together advances from Microsoft’s own model portfolio with a model architecture purpose-built for decision-making. Over time, this work could create opportunities to further explore how Microsoft models can be specialized for the high-volume decisions that sit inside applications and agentic systems.
For developers, Microsoft-Decision-1 expands the choices available when building applications and agents with Foundry. Rather than using the same model for every step, developers can combine generative and reasoning models with a decision model and choose the capability that best fits each part of the workflow. That model diversity is increasingly important as AI applications evolve from single-model experiences into systems that coordinate specialized models, tools, and actions.
How developers are building with Microsoft-Decision-1
Enterprise applications make decisions throughout a workflow: identifying a customer’s intent, selecting a tool, routing an incident or checking whether a response meets specified criteria. For example, a support application could classify a request as a billing question, a technical issue or an account-access problem. An agent could evaluate a proposed response against a quality rubric and determine whether to accept it, revise it or request review.
These decisions can sit alongside generative and reasoning steps in the same application. A language model might draft a response, while Microsoft-Decision-1 evaluates it against defined criteria. With Microsoft-Decision-1 in the Foundry catalog, developers can explore decision workloads through the platform they use to discover and deploy models across various scenarios:
-
Classification: Assign inputs to defined categories, such as customer intents or feedback themes.
-
Routing: Select a model, tool, destination or workflow from a set of options.
-
Evaluation and verification: Assess responses or inputs against specified criteria.
-
Workflow control: Decide whether to continue, retry, stop or send a case for review.
A different model for a different kind of workload
For developers, decision models introduce a different way to think about optimization. An AI application may make thousands of small choices around a comparatively small number of complex generative or reasoning tasks. Running every one of those steps through the same general-purpose model can overlook an opportunity to optimize the system around the work each step actually needs to perform.
That pattern is showing up across Microsoft. Xbox Research used Microsoft-Decision-1 to categorize more than 10,000 pieces of open-ended feedback into researcher-defined themes, reporting quality competitive with GPT-5 while running 80–100 times faster. Other Microsoft teams have also evaluated the model for assessing Copilot response quality and supporting adaptive replanning in Microsoft Discovery.
These scenarios matter because they are not simply classification benchmarks. They are examples of decisions embedded inside larger applications. As AI architectures become more modular, developers can consider generation, reasoning, and decision-making as distinct workloads, then select models based on what each step actually requires.
For benchmark methodology, measured results and additional examples, read the Command Line blog here.
Putting Microsoft-Decision-1 to work
For tasks with predefined outcomes, evaluate Microsoft-Decision-1 using representative inputs and the quality, latency and cost requirements of your application. Compare decision accuracy, including ambiguous cases, latency under expected traffic and cost per completed decision. Where confidence estimates inform automation or review thresholds, validate their calibration on your own data.
Start with a decision in your application: The following Python example shows how to compare support-routing decisions with expected labels and measure request latency. Replace the illustrative messages and category definitions with examples from your own workload.
Before running: Deploy Microsoft-Decision-1 in your Microsoft Foundry subscription and configure FOUNDRY_BASE_URL and FOUNDRY_API_KEY. Set FOUNDRY_MODEL to the model or deployment identifier specified in the quickstart. Keep credentials outside the source code. Model requests incur applicable usage charges.
Integration note: Confirm the route and authentication header before running it against your deployment.
#!/usr/bin/env python3 """Classify support requests with Microsoft-Decision-1.""" import json import os import urllib.error import urllib.request from collections.abc import Mapping from statistics import median from time import perf_counter from typing import Any from azure.identity import DefaultAzureCredential OPTIONS = { "billing": "Charges, invoices, refunds, or subscription payments", "technical": "Software errors, bugs, or integration failures", "account": "Sign-in, password, or account-access problems", } EXAMPLES = [ ("I was charged twice.", "billing"), ("The integration crashes during checkout.", "technical"), ("I cannot sign in after resetting my password.", "account"), ] AZURE_ENDPOINT = os.environ["AZURE_ENDPOINT"].rstrip("/") DEPLOYMENT_NAME = os.environ["DEPLOYMENT_NAME"] TIMEOUT_SECONDS = 60 TOKEN_SCOPE = "https://cognitiveservices.azure.com/.default" CREDENTIAL = DefaultAzureCredential() class DecisionAPIError(RuntimeError): """Raised when the API can't return a valid classification.""" def read_choice( payload: Any, options: Mapping[str, str], ) -> str: if not isinstance(payload, dict): raise DecisionAPIError("The API returned an invalid response.") answers = payload.get("answers") if not isinstance(answers, dict): raise DecisionAPIError("The API returned no answers object.") answer = answers.get("team") if not isinstance(answer, dict) or answer.get("type") != "choice": raise DecisionAPIError("The API returned an invalid answer.") choice = answer.get("choice") if not isinstance(choice, str) or choice not in options: raise DecisionAPIError( f"The API selected an unsupported team: {choice!r}." ) return choice def predict( text: str, options: Mapping[str, str], ) -> str: if not text.strip(): raise ValueError("text must not be empty.") if len(options) < 2: raise ValueError("options must contain at least two choices.") body = json.dumps( { "model": DEPLOYMENT_NAME, "state": text, "questions": { "team": { "type": "choice", "instructions": ( "Which team should handle this customer support " "request? Select exactly one team based on the " "primary problem." ), "criteria": dict(options), } }, } ).encode("utf-8") access_token = CREDENTIAL.get_token(TOKEN_SCOPE).token request = urllib.request.Request( f"{AZURE_ENDPOINT}/providers/microsoft/v1/systemone", data=body, headers={ "Authorization": f"Bearer {access_token}", "Content-Type": "application/json", "Accept": "application/json", }, method="POST", ) try: with urllib.request.urlopen( request, timeout=TIMEOUT_SECONDS, ) as response: response_body = response.read() except urllib.error.HTTPError as exc: raise DecisionAPIError( f"The API returned HTTP {exc.code}." ) from exc except urllib.error.URLError as exc: raise DecisionAPIError( f"Couldn't reach the API: {exc.reason}." ) from exc try: payload = json.loads(response_body) except (json.JSONDecodeError, UnicodeDecodeError) as exc: raise DecisionAPIError( "The API returned invalid JSON." ) from exc return read_choice(payload, options) def main() -> None: correct = 0 latencies_ms = [] for text, expected in EXAMPLES: start = perf_counter() predicted = predict(text, OPTIONS) elapsed_ms = (perf_counter() - start) * 1000 correct += int(predicted == expected) latencies_ms.append(elapsed_ms) print( f"Expected={expected}, predicted={predicted}, " f"latency={elapsed_ms:.1f} ms" ) accuracy = correct / len(EXAMPLES) print(f"Accuracy: {accuracy:.1%}") print( "Median request latency: " f"{median(latencies_ms):.1f} ms" ) if __name__ == "__main__": main()This small dataset illustrates the evaluation process. For meaningful results, use a larger, representative labeled dataset and inspect errors by category. The measured latency includes the client call and network time; sequential requests do not measure latency under production load. Assess cost and confidence calibration separately.
Your application defines the available options and resulting actions. Use evaluation results to determine where automation is appropriate and where additional validation or human review is needed.
Pricing
|
Model |
Deployment Type |
Input Tokens (#/1M) |
Output tokens (#/1M) |
|
Microsoft-Decision-1 |
US Datazone |
$0.042 |
N/A |
|
EU Datazone |
$0.042 |
N/A |
Get started
Explore Microsoft-Decision-1 in the Foundry catalog to make your first request. We look forward to seeing how developers use Microsoft-Decision-1 to bring structured decisions into their applications and agents through Microsoft Foundry.