Skip to content

TypeSafe Jev: what a decision-only model means for agents

TypeSafe Jev answers typed questions in about 100 milliseconds. What the model does, where it fails and what it means for agent workflows.

Jev, one model between a typed question and its answer: the question goes in with what it is asked against, a ticket, a diff, logs, and one of three typed answers comes out, a Choice, a Score or a Noul, each with its confidence.
On this page 5 sections

On September 15, TypeSafe AI came out of stealth with $40 million in seed funding led by DCVC and a first model called Jev. Since then it has shown up on almost every developer blog, even though it can neither chat nor write code. For teams working with AI agents it is still one of the more interesting releases of the year.

What Jev is

TypeSafe was founded by Diogo Almeida, who worked on RLHF and InstructGPT at OpenAI, the training methods behind ChatGPT. Jev is deliberately not a language model in the usual sense. You send it a state, for example a ticket, a log line or a shell command, together with a list of questions whose answer shapes you define up front. There are three kinds: a yes or no question (TypeSafe calls it a “Noul”), a pick from a fixed set of options (a “Choice”) and a position on a scale whose levels you describe yourself (a “Score”). What comes back is no text at all. Each question gets a probability distribution plus a confidence value.

Three cards showing the answer shapes Noul, Choice and Score with probabilities
The three question types. You define the possible answers up front and get a distribution plus confidence back.

In his detailed write up, Flavio Copes describes Jev as a kind of smart if statement. Ordinary code branches on values it can compute. As soon as the condition is a judgment, like whether a command is dangerous or whether a ticket contains enough information, you used to need a large language model or a regex that breaks on every edge case.

The numbers TypeSafe gives: most calls finish in around 100 milliseconds, input costs $0.042 per million tokens and output is free. All questions about one state are answered in parallel, so an extra question barely adds any latency. The model was trained with TypeSafe’s own method, Reinforcement Learning for Calibrated Decisions. The goal is that answers given with 90 percent probability turn out to be right about 90 percent of the time.

Bar comparison of time per decision and price per million input tokens on log axes
Time and price per decision compared on logarithmic axes. All values come from TypeSafe’s own measurements.

Where a model like this sits in an agent workflow

On the way from ticket to pull request, a coding agent makes lots of small decisions that have nothing to do with writing. Is the ticket clear enough to start? Does the task need a fast model or a strong one? Can this command run without asking anyone? Does the finished pull request match what the ticket asked for? Today these questions are usually answered by the same large model that writes the code, with the waiting time and cost that come with it.

Playbook run from ticket to pull request with steps, decisions and gates
One playbook run, simplified. Violet marks the steps where a language model writes, amber the decisions in between.

The first integrations show where this is heading. At launch, LangChain released a middleware that uses Jev to send each request to either a cheap or a powerful model, plus a second one that checks tool calls for risk before they run. At Browser Use, a team built an agent that finds a flight on Google Flights in just over seven seconds, because every click is a choice among numbered page elements and only typing text still needs a language model. Others have broken code reviews down into a risk matrix per file that Jev fills in.

In all of these examples the language model still writes the text or the code. Jev decides on the next step, while the loop, the safety rules and anything involving arithmetic stay in plain code.

What to keep a sober eye on

Jev is two weeks old. Direct access from TypeSafe still runs through a waitlist, while Vercel’s AI Gateway offers it without the wait. The headline figures of almost 200 times faster and more than 400 times cheaper come from TypeSafe’s own tests. TypeSafe itself places them at the high end of real world results. The accuracy comparison has a catch too: there is no independent ground truth, so what gets measured is agreement with the averaged judgment of two frontier models. Nobody outside the company has reproduced it yet.

The claim that Jev cannot hallucinate is narrower than it sounds. The model cannot return a value outside the options you defined, but it can still return the wrong allowed option. TypeSafe is also open about where Jev is weak: it does not count reliably, it cannot compare dates safely and it reads questions very literally. It does not understand images, audio or video yet.

What this means for teams

The real news is less about the model itself and more about what it says about how AI systems are built. Developer Hassan El Mghari had 1,018 research papers summarized and classified. The summaries from a language model cost $3.99, the classification with Jev cost 8 cents. His takeaway: different parts of a workflow will be handled by different models instead of one model doing everything.

For engineering teams this says a lot about how a good setup should be built. The work has to be broken into clear steps, otherwise nobody knows where a fast decision is enough and where a strong model needs to think. Every step needs a result that code can check. Models also need to be swappable, because the next new type of model is certain to arrive, probably sooner than expected.

If you want to try Jev, start with one unremarkable decision that a large model or a fragile rule makes today. Let Jev run alongside in shadow mode and log its answers. Then compare the confidence values with what was actually right. Automate the case where a mistake costs little first. Keep the questions and thresholds in one shared file, because that is the part a human needs to review.

At Kadmo, specialized agents work through playbooks in which every step ends with a gate, a clear verdict of pass, stop or pause. The model behind each agent can be swapped at any time. A structure like that lets you put new types of models such as Jev where they are strongest. If you want to see what this could look like for your team, we are happy to show you.

Sources

Keep reading

All articles

10x your engineering output.Keep the quality.

Live within days. Try it on one thing from your backlog, see the result, then decide.