When I talked about the Jev model half a month ago, I said one thing: for this kind of pure judgment and scoring routing model, big companies would implement it very quickly; there’s nothing to get hyped about at all.
But just two weeks later, Microsoft stepped in with Microsoft-Decision-1.
Microsoft didn’t even spend the money to acquire a startup team; it directly took the open-source Qwen3.5-9B and did post-training, specifically for single-forward-pass probability scoring over fixed options.
In the 36 evaluation benchmarks of JevBench, Microsoft’s model achieved an average accuracy of 83.5%, beating Jev’s 82.3%. Median latency was brought down to 85 milliseconds, about three times as fast as Jev, and 35 times faster than using GPT-6 Sol to run judgments.
Pricing is even more of a table-flip: only $0.042 per million input Tokens, and output Tokens are simply free.
This kind of model is essentially a pre-filtering node for Agents and complex systems. Microsoft itself uses it internally to clean up tens of thousands of player feedback items for the Xbox division and to validate actions at every step of automated workflows, boosting speed more than tenfold and cutting costs by two orders of magnitude.
What makes it absurd is that at the time, TypeSafe AI founder Diogo Almeida directly dismissed existing models as slow, expensive, and ruminating “System 2,” while Jev was “System 1.”
Looking back now at those first few days after Jev came out, the whole community’s cringeworthy hyping posture was such that one moment it was “generative AI is dead,” the next it was “a new paradigm of machine-native intelligence,” and with this rhetoric they even drove the valuation all the way up to $7.5 billion.
Now look: before it even lasted a month, this era-disrupting revolutionary technology was disrupted by Microsoft casually taking an open-source Qwen3.5-9B and tweaking it a bit. Accuracy overtaken, and nearly two times faster.
“Big Vs,” awkward, huh?
