Story
October 10, 2026
Jev’s breakout has turned routine AI decisions into a three-way war
TypeSafe sees Jev as a practical answer to AI’s gap between dazzling language and dependable automation; OpenAI has treated that traction as proof the category matters. An open-model challenge adds a sharper question for sensitive sectors: whether fast AI decisions should remain proprietary at all.
Jev arrived on Sept. 15 with a deliberately narrow pitch: not another chatbot, but a model built to make routine software judgments—sorting documents, routing customer messages and assigning scores—quickly and cheaply. TypeSafe says the model returns probabilities rather than prose, a design meant to spare businesses the cost of generating “a full written answer when all you need is a label.”1
The proposition caught on at startling speed. Within its first day on Vercel’s AI Gateway, Jev had been tried by nearly 13% of paid teams, more than twice the first-day reach of any prior launch there. Its core argument was economic: when applications make thousands of low-stakes calls, latency, price and error rates determine whether automation is viable at all.1
But TypeSafe’s own story also carries a warning about the limits of the pitch. Co-founder Diogo Almeida said building reliable decision models proved far tougher than expected: “Making the models reliable was so much harder than expected.”1 Developers must still set confidence thresholds and decide which uncertain calls go to human review.
On Oct. 6—three weeks after Jev’s release—OpenAI launched Decisions API, using its GPT-6 Luna model to produce structured, near-real-time choices. OpenAI’s Nikunj Handa said Jev had “inspir[ed] this whole thing”; the product had not been on the roadmap four weeks earlier.1 The response broadened the contest: Jev is purpose-trained for classification, while OpenAI is adapting a general model and adding image support.
TypeSafe then raised $870 million at a $7.5 billion valuation, saying a third of Fortune 500 companies were already using Jev—an extraordinary commercial vote of confidence for a model only weeks old.2
Yet the latest challenge is not merely from a larger closed rival. A widely shared post claimed that open-source pplx-decider v1.1 beat Jev on clinical decisions, scoring 643 of 669 correct against Jev’s 628, at 42% lower cost and similar speed. Its crucial selling point was governance as much as performance: hospitals could download the weights and run the model internally.
3