Microsoft built its decision model on Alibaba's weights

Published Oct 11, 2026

TLDR

Microsoft’s decision model is an open source Qwen3.5-9B that Microsoft post-trained into a closed product, priced at $0.042 per million input tokens with output free, and it landed nine days after five open-weight decision models shipped in the same week. It enables any team to buy routing and classification at 35x the speed of a general model, measured against benchmarks Microsoft chose and blind to its training. The difference between Microsoft’s version and the open nine is not the architecture, which is a single forward pass either way, it is the calibration: Microsoft sells probabilities it claims can be trusted at face value, and that claim, not the weights, is what the $0.042 buys.

Caption: the decision-model map as of October 2026: seven open-weight rows (green Apache, purple non-commercial) against two closed ones; every model shipped between September 30 and October 9.

Caption: the decision-model map as of October 2026: seven open-weight rows (green Apache, purple non-commercial) against two closed ones; every model shipped between September 30 and October 9.

The category arrived in one week

September 30 brought JEV-27B-VL, a 27.8B decision model with vision that collected 1.5 million downloads inside a week. Cloudflare shipped Clef and Clef-flash the next day, autotrust followed with GEV-26B-Decide, Torchcast put out a non-commercial 12B, Perplexity open-sourced its production router, and on October 9 Cloudflare completed the family with an omni variant that takes audio. Nine days, seven open-weight decision models, four of them Apache-2.0.

The category has a shared contract: hand the model a fixed set of options, get one probability per option in a single forward pass, with no generated text between question and answer. The economics follow from the shape: a one-pass 9B model answers in a handful of milliseconds on modest hardware, which is why the same contract also arrived from the edge side with Liquid AI’s d1 family, a 600M model with audio input that runs where no GPU goes.

Microsoft’s entry is a Qwen in a trench coat

The post-training base is the sharpest fact in the announcement: Microsoft built its first decision model on Alibaba’s open Qwen3.5-9B, the same family running production search at Ecosia as of October 9, and says the next versions will rebase on MAI and OpenAI models. A hyperscaler’s closed product is a Chinese open-weights model with a training run on top. The weights are closed, so nobody can check what changed between Qwen3.5-9B and Decision-1, which makes Microsoft’s benchmark claims the only public evidence of the delta: highest accuracy across 36 benchmarks with nearly 150,000 blind questions, 2.5x quicker than the runner-up, and decisions that hold on 98.7 percent of eight-way perturbations.

The pricing is the honest part: $0.042 per million input tokens, output free. At 300 input tokens per decision, a million decisions costs $12.60. The same million passes on a self-hosted 9B at 4-bit, about 5.2GB, run on a $400 mini-PC at battery-class watt draw and cost under a dollar in electricity. The category’s price floor is near zero on hardware you already own; what the hosted version sells is the SLA, the calibration, and the absence of an operational burden.

Caption: a million decisions: $12.60 hosted against under $1 in electricity self-hosted; the delta is the price of calibration with no ops.

Caption: a million decisions: $12.60 hosted against under $1 in electricity self-hosted; the delta is the price of calibration with no ops.

The clone exists, and it is one page of PyTorch

Hand a stock Qwen3-1.7B a question, mask every token in its vocabulary except the letters labeling your options, and one forward pass returns a probability for each. Nish Tahir’s October 10 walkthrough ships exactly that in about a page of PyTorch, and the Hacker News thread it hit (344 points) is the mainstream moment for the trick. What the clone does not give you shows up the first time you trust it: the scores are next-token probabilities, not the model’s belief that the answer is correct. A stated 90 percent that turns out right nine times out of ten has to be trained in. Not a free consequence of masking logits.

That boundary is the whole difference between the map’s green side and its red side. The open nine ship weights and let you run the mechanic anywhere; their confidence numbers are token probabilities until somebody pays for a calibration run. Microsoft’s product page sells the calibration directly: its probabilities are “part of the API,” usable to decide when to act, defer, or ask for review. For an agent harness gating money or actions on a score, that guarantee, priced at $0.042 per million, is the product the open weights do not yet include.

What to watch

  • Whether the open side ships calibrated variants: JEV’s team and Cloudflare both have the traffic to train it, and the first open model that publishes a reliability diagram turns Microsoft’s pricing power inside out.
  • The MAI rebase Microsoft promised: the Qwen base is the honesty anchor in the current version; an eventual MAI-base rebuild removes the one checkable lineage.
  • Whether GPT-6 Luna Decisions’ pricing moves under Decision-1’s $0.042: two closed vendors in one category of one-pass answers is a price war format.
  • Ecosia’s production numbers on open Qwen models: the same family Microsoft trained from is already serving search traffic at scale, per its own announcement, which shrinks the “closed is more reliable” argument to the calibration claim alone.
  • The clone’s calibration gap getting closed by the community: the HN thread’s 77 comments are the recruiting pool; a trained calibrated adapter on Qwen3.5-9B is a weekend-scale project for a lab with a labeled routing set.

Sources: Microsoft Command Line - Build your own decision model - HN thread - Marktechpost - TechNode on Ecosia - OpenAI Decisions API

Related on this site: pplx-decider-v1.1-27b - Clef - JEV-27B-VL - Torchcast Decision 12B - GEV-26B-Decide

Discussion

Be the first to comment

Start a discussion

Got a take on this, a rig to show off, or a benchmark that says otherwise? Sign up and start the thread - your comment publishes instantly once you're in.