Photo by Compagnons on Unsplash. Source: https://unsplash.com/photos/woman-working-on-computers-in-an-office-4G3vzMnC34M (Unsplash License).

Executive Summary

Microsoft shipped Microsoft-Decision-1 on Friday, a small model built for classification and routing rather than conversation. It runs on Alibaba’s Qwen3.5-9B, and Microsoft prices it at $0.042 per million input tokens with output tokens free. Microsoft claims it is more than 20 times cheaper than GPT-6 Sol for text classification. The company says it will rebase the model on its own weights and OpenAI’s, with no date.

The price is the story, not the benchmark. A decision model returns a label and a probability instead of a paragraph, so it can take over the cheap calls a general model currently handles at full price. More than 100 of these models are now listed, and the routing layer that sends easy requests to the cheapest model is where the saving actually sits. Treat the accuracy claim as a vendor number and test it on your own labels.

Every AI bill has a quiet line item, the requests that never needed a frontier model. A classification, a routing decision, a content flag, a short extraction. Teams pay chat-model prices for work a small model could do, because the routing logic is harder to build than a prompt.

Microsoft now sells that small model directly. Microsoft-Decision-1 went live on Friday through Microsoft Foundry, with OpenRouter to follow. It is a decision model, which means it returns a constrained, structured answer rather than free text. Microsoft pitches it for tasks where a wrong paragraph is worse than a wrong label.

The category is not Microsoft’s invention. TypeSafe AI shipped Jev in September and framed it as a model for machines, rated by probability rather than prose. Cloudflare, OpenAI, Liquid AI, Snowflake and Perplexity followed with their own. A public tracker now lists more than 100 decision models, which is what happens to a category that is cheap to enter.

The price is the product

Microsoft’s own numbers make the intent plain. Input tokens cost $0.042 per million and the output is free, because the answer is a handful of tokens. The company claims the model is 2.5 times faster than H2O-Lightning-4B and 2.8 times faster than Jev, and more than 20 times cheaper than GPT-6 Sol for text classification. Those are vendor measurements on vendor tests, and they should be read that way. The direction is what matters.

The economics are simple once the output is a label. A chat model bills for every word it produces, and a long reasoning trace costs real money. A decision model produces a fixed, tiny output, so the cost per decision collapses. The saving is not really in the model. It is in moving the easy half of your traffic onto a cheaper one.

Diagram comparing a general chat model with a decision model, and listing Microsoft-Decision-1 figures. 83.5 percent accuracy on 36 benchmarks, 92.2 percent confidence score, $0.042 per million input tokens and free output tokens, and a claimed 20 times lower cost than GPT-6 Sol for classification.
Where a decision model fits. It trades free text for a cheap structured answer, and the price is the pitch.

It runs on Alibaba’s weights, for now

The most interesting detail is the base model. Microsoft-Decision-1 is built on Qwen3.5-9B, from Alibaba Cloud. Microsoft says it will rebase on its own models and OpenAI’s, but gives no date. So the first version of Microsoft’s decision model is a tuned Qwen, which says as much about how fast this category moved as anything in the benchmark table.

The marketing line is that decision models do not hallucinate the way chat models do, because the answer space is constrained. That claim deserves scrutiny. A constrained answer can still be confidently wrong, and a model this cheap will be pointed at decisions nobody checks by hand.

What this does to your inference bill

The practical move is routing. Keep the general model for work that needs it. Send classification, extraction and simple decisions to the cheap one. The saving is largest where you already make millions of small calls, because the unit price is what scales. The risk is that routing adds a failure mode of its own. A classifier that misfiles a request sends it to the wrong place, quietly, at scale.

Ask three questions of your own stack. How many of your requests are a label rather than a paragraph. What would a twenty times lower unit cost do to that line of the bill. And do you have a labelled set you trust enough to test a cheap model against before you route to it.

Related reading. Our Enterprise AI Infrastructure report covers the four layers under the model and where they break, from hardware up to the development layer.

By Ivan Tarin

Ivan Tarin is a Principal Product Marketing Manager at SUSE, where he owns go-to-market strategy and positioning for a seven-product cloud-native portfolio spanning Kubernetes, virtualization, storage, security, and observability. A former full-stack developer who shipped production code for enterprise and public-sector clients including U.S. national laboratories, Ivan translates complex infrastructure and AI technology into messaging that lands with developers, platform teams, and enterprise buyers. He has presented at KubeCon, SUSECON, and AWS Developer Week, and is currently pursuing an MS in Artificial Intelligence at the University of Colorado Boulder.

One thought on “The Cheapest Inference on the Market Is a Classifier, Not a Chat Model”

Leave a Reply

Your email address will not be published. Required fields are marked *

Get the next one before it is old news

Independent analysis of cloud-native infrastructure, Kubernetes and data center economics. No vendor spin.