2026-10-09

Microsoft Launches Microsoft-Decision-1, a Qwen3.5-9B-Based Decision Model That Leads Average Accuracy on 36 Benchmarks at 85 ms Median Latency

AIModels🌍 North America

Microsoft launched Microsoft-Decision-1 on October 9, 2026, a model for fast decision-making available in Microsoft Foundry and through OpenRouter. It is post-trained from Qwen3.5-9B and returns a calibrated probability for each answer option in a single pass, for yes/no, multiple-choice and rating questions as well as rubric grading of AI responses and agent actions. Microsoft aims it at routing, classification, prioritization, verification and workflow control, and claims the highest accuracy in its comparison of 36 benchmarks and 147,137 questions, with the fastest measured response: an 85 ms median (p50), about 35x quicker than GPT-6 Sol. Satya Nadella's announcement thread says Microsoft is already testing it "from incident response and quality control to scientific discovery".

What a decision model does

Where an LLM generates text or reasons at length, a decision model produces a structured output that software can act on immediately, at very low cost. Microsoft-Decision-1 takes a fixed set of options and scores each one through a simple structured API call. It will soon be rebased on other models, including Microsoft AI (MAI) and OpenAI models. This blog covered a similar launch in Liquid AI's open d1 models two days earlier.

Microsoft frames the design around five challenges:

  • Speed. Adding 100 ms to each of 20 sequential decisions adds two seconds.
  • Generalization. Evaluation spans benchmarks kept blind from training: routing, ranking, long context, multilingual, out-of-distribution, reasoning and safety. Microsoft also took several top public models from the JevBench leaderboard and ran them on 36 additional public and private benchmarks.
  • Robustness. Perturbing the same request eight ways flips the decision on 1.3% of perturbations on average, and on none when option descriptions are paraphrased or options are reversed or shuffled.
  • Calibration. Probabilities are part of the API, so a 90% prediction should be right about 9 times in 10.
  • Safety. On 5,250 requests across 11 benchmarks covering harmful content, jailbreaks and prompt injection, the model refused harmful behavior while keeping high utility.

Benchmarks

Average accuracy over the 36 benchmarks puts Microsoft-Decision-1 first at 83.5. Calibration (100 means confidence exactly matches how often the model is right) puts it third at 92.2.

Bar chart of average accuracy over 36 benchmarks: Microsoft-Decision-1 at 83.5, ahead of Jev 1.13.0 at 82.3 and Quyet-1.0-Large at 81.9. From Microsoft's launch post and Nadella's thread.

Bar chart of calibration scores over 36 benchmarks: Jev 1.13.0 at 93.7, Quyet-1.0-Large at 93.1, Microsoft-Decision-1 at 92.2. From Microsoft's launch post and Nadella's thread.

ModelAccuracyCalibrationMedian latency
Microsoft-Decision-183.592.285 ms
Jev 1.13.082.393.7219 ms
Quyet-1.0-Large81.993.1380 ms
Surogate Rune 26B-A4B79.791.8380 ms
GPT-6 Luna Decisions79.489.9299 ms
deck-31B77.883.5394 ms
H2O-Lightning-4B (v1.1 for latency)77.291.8201 ms
Strands-Decider 2B54.8*n/an/a
GPT-6 Soln/an/a2.68 s

Strands-Decider 2B (AWS) was scored on 23 of the 36 benchmarks, so its average covers only those. Latencies are JevBench adjusted medians, with Microsoft-Decision-1 measured through Foundry in the same region; its p95 is 125 ms.

The accuracy lead is modest: 1.2 points over Jev 1.13.0 and 1.6 over Quyet-1.0-Large, which holds first place on JevBench. On calibration the model trails Jev by 1.5 points and Quyet-1.0-Large by 0.9, while staying ahead of the other four models scored; every model shown is within 17 points of perfect.

Speed is the widest gap. Against Quyet-1.0-Large the model is 4.5x quicker (380 / 85), and against H2O-Lightning-4B v1.1, the fastest of the other decision models, 2.4x (201 / 85). Against GPT-6 Sol, 2.68 s over 85 ms is about 31x with the rounded figures in the table; Microsoft puts it at about 35x.

Inside Microsoft

Microsoft's own teams have been testing the model:

  • XBOX Research sorted more than 10,000 open-ended feedback items from surveys, Steam and Twitter/X into fixed themes. Quality was competitive with GPT-6 Sol, at over 14x the speed and 200x lower cost.
  • The Copilot team, scoring the quality of chat and agentic responses, found it competitive with GPT-5.6 Luna and 100x faster.
  • Incident response: for on-call engineers, it retrieved knowledge better and faster than an LLM.
  • Microsoft Discovery: scores were 46x more consistent than an LLM-based score, at three times the speed, and adaptive replanning ran nearly four times faster.

Two demos make the same point: a classification run against GPT-6 Sol, and a computer-use agent that buys a backpack faster than one driven by GPT-6 Sol.

Where to use it

Microsoft's list of use cases includes:

  • agent controls (continue, stop, retry, hand off)
  • model routing
  • AI judging and data labeling
  • incident response routing
  • safety and security screening
  • computer and UI use
  • robotics
  • scientific discovery

Context

GPT-6 Sol, the latency reference above, launched on September 22 alongside Luna; see GPT-6 Sol and Luna. The planned rebase onto MAI models follows Microsoft's Windows hybrid-intelligence announcement, which put the MAI-Code-1.1 Flash coding model on the PC.

Read next