Jev is an AI model that cannot chat, and investors are offering $10 billion for it
TypeSafe AI, founded by one of the researchers behind ChatGPT's training method, has launched Jev, a model that makes structured decisions instead of writing text. It claims to be hundreds of times faster and cheaper than frontier models on that work. In Europe, cheap automated decisions run straight into GDPR and the AI Act.
TypeSafe AI, a San Francisco start-up, opened early access to Jev on September 15. Jev is a model that does not write. You give it a situation and a question with a fixed set of possible answers, and it returns one of those answers with a probability attached. Its launch video drew around 40 million views on X in its first week, and the Financial Times reports investor offers valuing the company at more than $10 billion, against $200 million at its seed round. That is a lot of money for a model that cannot hold a conversation. It also makes sense: most of what businesses want to automate is a decision.
A model built to answer one question at a time
TypeSafe’s chief executive is Diogo Almeida, a former OpenAI researcher who worked on the training method behind ChatGPT. He founded the company with Erik Gafni and Sasha Sheng. Almeida’s argument, as he put it to TechCrunch, is that today’s models are optimised for human language, which is not what software needs.
In practice, a developer sends Jev structured data, such as a customer’s recent messages and transactions, and asks something like whether the customer is asking for a refund. Jev answers yes or no, or picks from up to 255 options, and says how confident it is. It generates all of its outputs at once instead of word by word, which is where the speed comes from. Because it can only choose from the options it is given, it cannot invent an answer outside them, which is what TypeSafe means when it says Jev does not hallucinate.
The numbers are TypeSafe’s own: responses in 70 to 500 milliseconds where conversational models can take seconds or minutes, input at $0.042 per million tokens with output free, and workflows up to 193.6 times faster than frontier models on its published benchmarks. TechCrunch reports early independent results that are smaller but still large. Vercel found Jev 5 to 18 times faster than Luna 5.6, OpenAI’s high-volume model, at classifying content for safety, with better accuracy, and the start-up Bryo found it 10 to 20 times cheaper than Gemini.
Why the hype is partly deserved
The interesting claim is calibration. TypeSafe trains Jev with a method it calls reinforcement learning for calibrated decisions, designed so that when Jev says it is 90% sure, it is right about 90% of the time. Language models are bad at this. Ask one how confident it is and you get a number that sounds plausible and means little.
A calibrated probability is what automation needs. It lets a company say: act automatically above 95%, send everything else to a person. Armin Ronacher, a well-known developer, made the obvious point to TechCrunch that an answer at 50% is a coin toss. The value is that Jev tells you when it is tossing a coin.
There are good reasons for caution. The headline numbers come from workflows TypeSafe’s own team selected, which the company acknowledges could bias the results. Jev handles text and structured data but not images yet, and it is no use for open-ended reasoning or writing, as Tom’s Hardware notes. “No hallucinations” also does not mean no mistakes. A model that picks the wrong option from a fixed list with 92% confidence is still wrong.
My view: Jev is the most practical idea in AI this autumn, and $10 billion is a bet that it works outside TypeSafe’s own test set. Those are two different claims. The first one looks right. The second will take months of customers’ own data to settle.
In Europe, a decision is a regulated thing
This is where the European angle matters. A model whose purpose is to make decisions cheaply invites companies to make more of them automatically, and EU law has views on that.
Under Article 22 of the GDPR, people have the right not to be subject to a decision based solely on automated processing when it has legal or similarly significant effects on them. A refund request is probably fine. Turning down a loan, an insurance claim or a job applicant is not, unless the decision is necessary for a contract, allowed by law or based on explicit consent, and even then the person can demand human review and contest the result.
The AI Act adds another layer. Systems used for hiring, creditworthiness and access to essential services are high-risk, with rules that apply from December 2027: risk management, data governance, logging, human oversight. A Jev-based system screening job applications would fall in that category regardless of how cheap each decision is.
There is also the practical question of where the data goes. TypeSafe says the service runs from the US West Coast. A European company sending customer records to it needs a lawful basis for the transfer and a data processing agreement, the same as with any American provider.
What a business should do with this
Look for the decisions you already pay people or language models to make. Sorting incoming email, routing support tickets, flagging invoices that do not match an order, checking whether a document is complete: these are high-volume, low-stakes choices where a fast and cheap model with honest confidence scores could change the economics. If you are already using a model like GPT for this, test Jev against it on your own data.
Set thresholds before you go live. The confidence score is only useful if you decide in advance what happens at each level, and who checks the uncertain cases. Write that down. It is also what a regulator will ask for.
Keep consequential decisions with a person. For anything that affects someone’s money, job or access to a service, use Jev to sort and prepare the case for the person who decides. That is good practice anyway, and in the EU it is the difference between a compliant system and a GDPR complaint.
And wait for the second round of evidence. Jev is two weeks old, its benchmarks are mostly its own, and early access is still limited. Watch what the developers testing it publish over the next couple of months before building anything you cannot easily switch away from.


