Microsoft releases Decision-1 routing and classification model
Microsoft releases Decision-1, a Qwen3.5-9B decision-scoring model for routing and classification tasks, available in Foundry and OpenRouter.

What changed
Microsoft has released Microsoft-Decision-1, a decision-scoring model post-trained from Alibaba’s Qwen3.5-9B. The announcement came on October 9, 2026. The model is now available in Microsoft Foundry and through OpenRouter.
Unlike general-purpose language models, Microsoft-Decision-1 returns calibrated probabilities for each fixed answer option instead of generating text. It supports yes/no questions, multiple-choice, ratings, classification tasks, and rubric-based grading of AI responses. The model has a 32,768-token context window and runs as a hosted API only through Microsoft Foundry (measured on Azure) and OpenRouter. There are no open weights or quantized variants available.
According to Microsoft’s own measurements across 36 benchmarks with 147,137 questions, Microsoft-Decision-1 achieved 83.5% average accuracy. Latency measured at p50 is 85 milliseconds, with p95 at 125 milliseconds. The model is priced at $0.042 per million input tokens with free output.
Why it matters
Decision models represent a narrowing of what a language model does, built specifically for outputs that software can act on immediately. For enterprise workflows that need routing, classification, and agent control decisions, this specialization addresses real constraints in production systems.
The speed advantage is material. Microsoft reports the model runs 4.5 times quicker than Quyet-1.0-Large and 35 times quicker than GPT-6 Sol, which took 3.01 seconds per request. For high-volume classification tasks, latency compounds across thousands of decisions daily. Reduced latency directly reduces infrastructure costs and improves application responsiveness.
The cost profile also matters for scale. At $0.042 per million input tokens against $0.10 per million for GPT-6 Luna Decisions, the per-token expense is lower. For teams running millions of routing decisions annually, this pricing differential adds up.
Microsoft’s internal results suggest concrete use cases. In Xbox Research, the model sorted 10,000 plus feedback items 14 times faster and 200 times cheaper than GPT-6 Sol. For Copilot quality control, it was competitive with GPT5.6 Luna while running 100 times faster. In scientific discovery workflows, it was 46 times more consistent than LLM scoring and 3 times faster.
The calibration metric matters too. Microsoft reports calibration at 92.2 out of 100 (where 100 is perfect), second only to Quyet-1.0-Large at 93.1. Well-calibrated probabilities let developers set thresholds and confidence cutoffs based on the actual likelihood of correctness.
However, the model card explicitly restricts use: it is not for text generation, open-ended Q&A, chat, translation or summarization. It cannot handle images, audio or video. Outputs contain no explanations or rationales, only JSON probability scores. Microsoft explicitly states it is not for sole automated decisions on credit, employment, housing, healthcare or legal rights.
What to test
Before adopting Microsoft-Decision-1, teams should verify several vendor claims and assumptions.
Accuracy and fairness across domains. The 36 benchmarks are Microsoft’s own. The composition, difficulty distribution, and domain mix are not published. Test the model against your specific classification and routing tasks to see if the 83.5% average translates to your use cases. Check performance across demographic groups and edge cases relevant to your application.
Latency under production load. The 85 ms p50 latency is measured through Microsoft Foundry. Real-world latency depends on network hops, concurrent request load, and token count. Measure end-to-end latency from your application to the API under realistic traffic patterns.
Calibration in practice. While Microsoft reports a calibration score of 92.2, test whether the model’s confidence scores align with actual correctness in your domain. Set a confidence threshold and measure the precision and recall at that threshold to validate the probabilities are usable for your escalation logic.
Comparison with alternatives. The benchmark results show Microsoft-Decision-1 ahead of competitors, but measurement methods differ. H2O.ai disputes the latency comparison in Microsoft’s chart, claiming its own measured latency is 29 ms, not the 210 ms shown in the chart. Run your own comparative tests against Quyet-1.0-Large and H2O-Lightning-4B if those are options for your use case.
Robustness to input variation. Microsoft reports the model flipped decisions on 1.3% of perturbed requests. Test how the model performs when your input data is paraphrased, contains typos, or uses terminology outside the training distribution.
Cost at scale. Calculate your expected token consumption for your actual routing and classification volume. Ensure the $0.042 per million input tokens pricing fits your budget when multiplied by your annual decision volume.
API stability and versioning. OpenRouter notes that model weights update continually while the API shape stays fixed. Verify how updates are communicated and whether they affect your thresholds or escalation logic.
The conclusion
Microsoft-Decision-1 fills a gap: fast, cheap decision-scoring for fixed option sets. For teams running high-volume classification and routing workloads where latency and cost matter, the specialization offers clear benefits. The calibrated probabilities give developers a way to set confidence thresholds and escalation rules.
The tradeoffs are clear. You lose explainability, multimodal input, and the flexibility of general-purpose language models. You gain speed and cost efficiency. The model is not for open-ended tasks or sole automated decisions on sensitive categories.
The real question is whether it outperforms a smaller, quantized open-weight model fine-tuned on your specific classification task and self-hosted. The benchmarks are vendor-run and the latency comparison methodology differs from competitors’ reported figures. Both of those warrant skepticism in procurement decisions.
What to watch: whether Microsoft rebases future versions on its own models or OpenAI models as planned, how third-party benchmark boards rank it, and whether enterprises report comparable latency and accuracy in production deployments.
AI Tool Herald may earn a commission from some links on this site. It never changes what we report or recommend. Affiliate disclosure
Related stories

Claude Dynamic Workflows: 1,000 Parallel Agents Now Available
Anthropic added dynamic workflows to Claude Managed Agents, enabling up to 1,000 AI agents to run in parallel per execution through managed agent infrastructure.

Nace AI Open-Sources Drex 1.5 Decision Model for Option Scoring
Nace AI open-sources Drex 1.5, a 9B decision model that scores multiple options in one forward pass, ranking first under 10B parameters on Decision Index.

Cloudflare releases Clef-omni with audio and video support
Cloudflare launches Clef-omni, a decision model processing audio, video, images, and text in one API call, with price cuts and speed gains for existing models.