Skip to content
Technology

Microsoft Beats Jev's Benchmarks a Day Later

Microsoft matched Jev's price and topped its benchmark score one day after TypeSafe closed a $7.5 billion round built on being first.

6 min read · 1,310 words · 7 sources
Two sprinters neck and neck on a track, no gap between them
Two sprinters neck and neck on a track, no gap between them. Photo · Pexels
“TypeSafe spent two years building Jev, then watched it become a $7.5 billion company in weeks. Microsoft needed one day to match its price and beat its benchmark score, using a model it built on someone else's open weights. Does being first in artificial intelligence survive getting copied this fast?”

Jev skips writing altogether. A customer sends it a block of facts and a short list of fixed choices. Jev picks one, with a confidence score attached, in under a second2. TypeSafe, the company that built it, calls this a decision model. It is a model trained to choose among options, instead of composing a sentence. Think of the difference between a checkbox and an essay2. A support team might ask Jev whether a refund request fits one of six policy categories. It answers with a category and a confidence number, past any paragraph explaining its reasoning2. On Sept. 15, Jev shipped to limited early access. By Oct. 9, TypeSafe had turned that single model into a $7.5 billion company1. That same evening, Microsoft shipped a rival. It beat Jev's own benchmark score, at Jev's own price34.

A Model Built From Scratch, Priced at $7.5 Billion

TypeSafe spent two years in stealth before Jev's launch2. Diogo Almeida, the company's cofounder and CEO, built it on a simple bet. Four years of large language models had made computers fluent in human language, he argued. Fluent language, in his view, is the wrong tool for most software decisions. "We have been super good at human language for four years, but it's not useful for automation because computers speak a different language," Almeida said1. Jev crossed one million users within days of launch, by TypeSafe's own count. About a third of Fortune 500 companies now run it somewhere in their stack12. On Oct. 9, Andreessen Horowitz led an $870 million round. That round priced the company at $7.5 billion12. General partner Martin Casado joined its board12. Almeida frames the gap between Jev and a general chatbot as a difference in kind, beyond degree. "Jev achieves similar intelligence to existing large language models on System One tasks while running two orders of magnitude faster," he said2. TypeSafe named that category System One, after the fast, intuitive thinking the psychologist Daniel Kahneman once studied2.

Chart 1

Decision-1 answers in 85 milliseconds against Jev's 240

Median response time on the JevBench leaderboard, milliseconds, Oct. 9, 2026

Each column is one model's typical answer time; the shorter magenta column, Decision-1, responds nearly three times faster than Jev.

Source: cellcog.ai [4]. Chart by The AILately.com Desk.

The numbers behind this chart
ItemValue
Jev240 ms
Decision-185 ms

Two Rivals Answer Within Three Weeks

The category Jev opened stayed open briefly. OpenAI previewed its own answer on Sept. 30. The company called it the Decisions API, and built it on its existing GPT-6 Luna model, rather than training something new7. Sam Altman explained the trade-off directly. "By focusing the model on that choice, we can make it extremely fast while keeping capabilities like image understanding, broad language support, and safety protections," he said7. The Decisions API reached general release at OpenAI's developer conference on Oct. 6. That landed three weeks after Jev's debut6. Nikunj Handa, who leads the product at OpenAI, credited Jev directly for the pace. The feature, he said, had yet to reach OpenAI's roadmap a month before it shipped6. A model meant to prove a new category now had two frontier labs building inside that same category within a month. Both OpenAI and Microsoft had the budgets to copy an idea fast, once TypeSafe proved the idea sold.

Microsoft Matches the Price, Then Beats the Score

Redmond moved last and fastest. Its new Decision-1 model reached public preview on Oct. 9, the same day TypeSafe's round closed3. Jev is TypeSafe's own architecture, trained end to end. Decision-1 is a tuned version of Alibaba's open-weight Qwen3.5-9B, adapted by Microsoft's own engineers rather than trained from a blank page3. Both models now charge the identical price. Four cents per million input tokens, with output free on each34. Achint Srivastava, a Microsoft vice president, framed the pitch in plain terms. "Decision models are purpose-built to deliver structured outputs that software can immediately act on," he wrote in the announcement3. An outside tally on the JevBench leaderboard, the public scoreboard both companies now get measured against, ran the two head to head. Decision-1 answered in 85 milliseconds against Jev's 2404. It scored 83.5 percent accuracy across 36 benchmarks, against Jev's 82.34. Jev still leads on one measure that matters most when a decision skips a human's review. Its confidence scores track its real accuracy more closely, at 93.7 against Decision-1's 92.24.

Chart 2

Three rivals answered Jev within 25 days of its launch

Decision-model launches and funding events, Sept. 15 to Oct. 9, 2026

Each mark is one dated event; the magenta marks are rival moves, the gray marks are TypeSafe's own milestones.

Sources: TechCrunch [1]; Unite.AI [2]; Command Line (Microsoft) [3]; Fortune [6]; TechCrunch [7]. Chart by The AILately.com Desk.

The numbers behind this chart
DateEvent
Sept. 15, 2026Jev launches in early access
Sept. 30, 2026OpenAI previews the Decisions API
Oct. 6, 2026OpenAI's Decisions API goes live
Oct. 9, 2026TypeSafe closes its $7.5 billion round
Oct. 9, 2026Microsoft ships Decision-1

The Case That Jev Repeats Old Technology

Some observers reject the premise that a new model category exists at all. Anastasios Angelopoulos runs Arena, a platform that evaluates AI models for a living. He said so three weeks before Microsoft proved his point for him. "It's unclear to me what makes these models different from standard 'zero-shot classifiers', which are relatively well-known technology," he told the Financial Times5. A zero-shot classifier sorts an input into one of several categories. It skips any training on examples of that exact task first, a technique that predates Jev by years. Angelopoulos's argument, read against Microsoft's launch, gains weight fast. Suppose a decision model really is a known technique wearing a new name. Then the fastest path to competing with one is simply building one. That is exactly what Microsoft did, in a single day, using weights someone else trained first.

What Each Side Would Need to Be True

James Hardiman, a DCVC general partner and Jev investor, argues the benchmarks miss the point. "There's this kind of intangible element I think to some of these things that the benchmarks don't capture, and that's why I think actually in production, Jev will perform better than these kind of fast follow clones that we started to see come out," he said6. For Hardiman's case to hold, Jev needs an edge a public leaderboard misses. That edge could be the fine details of how it was trained. It could also be the trust TypeSafe already built with a third of the Fortune 500, who adopted Jev before any rival existed. Angelopoulos's case, by contrast, needs Microsoft's one-day answer to keep working at scale. It needs to hold up past any benchmark built to be legible to reporters. Only one of them gets to be right. A model category less than a month old is already testing which one. Anyone weighing a vendor contract this quarter can skip waiting for that answer to settle everywhere. They need to know whether their own workload rewards TypeSafe's purpose-built training. Or whether a cheaper, fast-follow model does the job just as well.

By the numbers

  • $7.5 billion: TypeSafe's valuation after its Oct. 9 funding round, up from roughly $200 million weeks earlier12.
  • 85 milliseconds: Decision-1's median answer time on the JevBench leaderboard, against Jev's 2404.
  • One million: users Jev crossed within days of its Sept. 15 launch, by TypeSafe's own count2.
  • Three weeks: the gap between Jev's debut and OpenAI's Decisions API reaching general release67.
  • $0.042: what both Decision-1 and Jev charge per million input tokens, with output free on both34.
  • 36: the benchmarks, spanning nearly 150,000 questions, where Microsoft says Decision-1 topped the field34.
  • A third: the share of Fortune 500 companies TypeSafe says already run Jev somewhere in their stack1.
  • 93.7 against 92.2: Jev's calibration score against Decision-1's, the one benchmark where TypeSafe's model still leads4.
  • Two years: how long TypeSafe spent building Jev in stealth before its Sept. 15 launch2.

What to watch

Watch whether TypeSafe answers with an audit of its own. An outside leaderboard has set the terms of this fight in public now. Decision-1's calibration score deserves a second look too, as it reaches more customers. A gap that looks small in a lab can widen once a system runs unattended in production. OpenAI's Decisions API runs on an existing model, past a purpose-built one. It may still settle the question Angelopoulos raised, about whether a decision model ever needed its own architecture. The next 30 days of enterprise contracts will answer that, past any leaderboard update. That answer will show which investor read this market correctly.

Sources

  1. Marina Temkin, "The maker of non-text AI model Jev valued at $7.5B just weeks after launch," TechCrunch, Oct. 9, 2026, https://techcrunch.com/2026/10/09/the-maker-of-non-text-ai-model-jev-valued-at-7-5b-just-weeks-after-launch
  2. Evan Mercer, "TypeSafe AI Raises $870M Series A at $7.5B Valuation to Ship More AI Models," Unite.AI, Oct. 9, 2026, https://www.unite.ai/typesafe-ai-raises-870m-series-a-at-7-5b-valuation-to-ship-more-ai-models/
  3. Achint Srivastava, "Introducing Microsoft-Decision-1, our model for fast decision-making," Command Line (Microsoft), Oct. 9, 2026, https://commandline.microsoft.com/microsoft-decision-1-model-foundry/
  4. cellcog.ai, "Microsoft-Decision-1: Price, Benchmarks, vs Jev," cellcog.ai, Oct. 9, 2026, https://cellcog.ai/blog/microsoft-decision-1/
  5. Nicole Jeffrey, "Start-up valued at $200mn fields $10bn offers after Jev launch," Financial Times, Sept. 25, 2026, https://ontimebrief.com/en/2026/09/25/start-up-valued-at-200mn-fields-10bn-offers-after-jev-launch/
  6. Wen Shao, "Jev, an AI for making quick decisions, has been a viral hit. But OpenAI is hot on its heels," Fortune, Oct. 8, 2026, https://fortune.com/2026/10/08/jev-an-ai-for-making-quick-decisions-has-been-a-viral-hit-in-silicon-valley-but-openai-is-hot-on-its-heels/
  7. Tim Fernholz, "OpenAI's Jev clone could help the frontier lab stop its swarming agents," TechCrunch, Sept. 30, 2026, https://techcrunch.com/2026/09/30/openais-jev-clone-could-help-the-frontier-lab-stop-its-swarming-agents/

Cite this piece

Ryan Elliott Dennis, "Microsoft Beats Jev's Benchmarks a Day Later," AI Lately, Oct 11, 2026, https://ailately.com/articles/microsoft-beats-jev-benchmarks

Tags: TypeSafe AI · Jev · Microsoft · decision models · OpenAI · Decisions API

Related, lately

More

Keyboard

j k
Move through a list
Enter
Open the selected piece
/
Search, and browse every column
g then a
Articles
g then s
The Signal
g then o
Opinion
g then b
Analysis
g then p
People
?
This sheet
Esc
Close