Microsoft-Decision-1 bills input only, at $0.042 per million
Microsoft-Decision-1 charges nothing for output tokens and $0.042 per million input tokens. That pricing fits what it returns: no prose, only a probability for each answer you let it choose from.
Microsoft described the model in a post on the Microsoft blog. It is live in Microsoft Foundry now and is coming soon through OpenRouter, pitched for routing, classification, prioritization, verification and workflow control inside existing apps and agents.
At a glance
- Microsoft took Alibaba's open-weight Qwen3.5-9B and post-trained it for single-pass decision scoring, and says it will later rebase the model on other foundations, including MAI and OpenAI models.
- In Microsoft's 36-benchmark comparison of nearly 150,000 questions it posted the highest accuracy and ran 4.5 times quicker than runner-up Quyet-1.0-Large and 35 times quicker than GPT-6 Sol.
- Every benchmark, speed figure and internal case study comes from Microsoft itself, and the post gives no per-call cost for the rival models it says Decision-1 beats.
If you have not been following the Qwen line: Qwen3.5-9B is a dense 9-billion-parameter model released under Apache-2.0 with downloadable weights. According to AI/TLDR, it came out on 2 March 2026 alongside 4B, 2B and 0.8B variants, after the flagship Qwen3.5-397B-A17B arrived on 16 February 2026. Microsoft Foundry has carried Qwen models before, though Qwen3-32B in its catalog is supported only for fine-tuning, with no base model inference.
Output tokens are free, and P50 latency is about 35 times lower than GPT-6 Sol's
The price list has two lines: input at $0.042 per million tokens, output free. For a sense of scale, AI/TLDR lists third-party hosted pricing for the much larger open-weight Qwen3.5-397B-A17B (FP8) on DeepInfra at $0.54 per million input tokens and $3.40 per million output tokens.
Microsoft calls speed the first problem it had to solve. Each decision adds delay, especially when one step waits on another: by the company's arithmetic, adding 100 milliseconds to each of 20 sequential decisions adds two seconds to the workflow. In its benchmarking, Decision-1 was the fastest model measured, 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, with P50 latency about 35 times faster than GPT-6 Sol.
Decision-1 led a 36-benchmark comparison spanning nearly 150,000 questions
Microsoft says the model achieved the highest accuracy in its 36-benchmark comparison, with nearly 150,000 questions drawn from benchmarks kept blind from training. The company frames the point as generalization: a model tuned for one benchmark can look strong and still fail on tasks it never saw, so the blinded sets span routing, ranking, long context, multilingual and out-of-distribution tasks, reasoning and safety.
Microsoft also took several top public models from the open leaderboard JevBench and ran them across 36 additional public and private benchmarks, where Decision-1 again performed best. The appendix names the models benchmarked: Quyet-1.0-Large, Surogate Rune 26B-A4B, OpenAI's GPT-6 Luna Decisions, deck-31B, H2O-Lightning-4B and Strands-Decider 2B.
Shuffling or reversing the options flipped zero decisions
Equivalent inputs should produce equivalent decisions, and Microsoft tests this by perturbing the same request in eight ways. The perturbations mirror production mess: paraphrased states and instructions, rewritten option descriptions, reordered choices, changed keys and harmless formatting noise. Decision-1 changed its decision on 1.3% of perturbations on average, with zero flips when option descriptions were paraphrased or options were reversed or shuffled.
On safety, Microsoft ran 5,250 requests across 11 benchmarks covering harmful content, jailbreak attempts and prompt injection. The company says the model refused harmful behavior while retaining a high degree of utility, meaning it did not block harmless requests needlessly. No refusal rate or false-positive rate is published alongside that claim.
Xbox Research sorted more than 10,000 reviews at 200 times lower cost
Testing Catalog quoted Microsoft on how widely the model is already in use:
We're already testing it across Microsoft for everything from incident response and quality control to scientific discovery.
Xbox Research sorted more than 10,000 open-ended pieces of feedback and reviews from surveys, Steam and Twitter/X into a fixed set of themes chosen by researchers. Decision-1 was competitive on quality with GPT-6 Sol while running over 14 times faster and 200 times less expensive. The Copilot team, which measures chat and agentic response quality, found it competitive with GPT5.6 Luna and 100 times faster.
For on-call engineers pulling knowledge from logs, tickets, calls and messages during live incidents, it performed better and faster than an LLM. In Microsoft Discovery's adaptive replanning, Decision-1 scored as 46 times more consistent than the LLM-based score at three times the speed, and made adaptive replanning nearly four times faster.
The model returns one calibrated probability per option in a fixed set
Decision-1 does not compose an answer. You send a fixed set of options through a structured API call, and in a single pass it returns a calibrated probability score for each one. Supported formats are yes/no, multiple choice and ratings, plus rubric-based grading of AI responses and agent actions. Think of an essay exam versus an answer sheet where you shade one bubble: the sheet is quicker to fill in and trivial for a machine to read.
Microsoft stresses that the probability is part of the API, not only a ranking score. Applications use it to decide when to act, defer or ask for review, so a 90% prediction should be right about nine times out of 10 on representative cases. The base Qwen3.5-9B interleaves Gated DeltaNet linear attention, roughly three quarters of its layers, with conventional softmax attention, and has a native context of 262,144 tokens.
All of the evidence here is Microsoft's own: its benchmark selection, its latency measurements, its internal teams. The post gives no absolute latency in milliseconds and no cost basis behind the 200-times figure, so you cannot check that claim against your own GPT-6 Sol bill. In our view, free output tokens are a smaller gift than the headline suggests, since a model that scores a fixed option list produces very little output by design.
OpenRouter and the MAI rebase. Two things are promised without dates. OpenRouter access is listed only as coming soon, and the rebase onto other models, including MAI and OpenAI ones, has no timeline. Microsoft also says it will keep shipping updates that use new evaluations and data to improve quality, confidence and cost. For now, the only version you can test is the one built on Qwen3.5-9B, in Foundry.
Related stories
- OpenAI's Decisions API beats its own engine by up to 10x
- AWS's Strands Decider 2B picks from your list, never writes
- Claude Max subscribers get up to $200 a month for the API
- Short prompts to Claude Haiku 5.5 cost 90% less
- Google's Nano Banana 2.1 costs $0.076 per 4K image
- Microsoft brings streaming MAI-Transcribe-2 to its APIs
Comments
No comments yet. Be the first.
Join the conversation
Sign in with Google to leave a comment. Your name and avatar come from your Google profile, and the comment appears after moderation.
