TypeSafe AI released Jev on 15 September 2026. The founder, Diogo Almeida, worked on RLHF at OpenAI. The pitch is a “System One model”: it does not generate text. You send it a block of text (the “state”) and a list of typed questions, and it returns an answer to each question with a probability distribution, all in one forward pass. Three question types exist: Choice (pick one of up to 255 options), Score (a level on a rubric) and Noul (a yes/no probability). Within two weeks the launch produced a wave of “universal classifier”, “40x faster”, “cannot hallucinate” headlines. Here is what each claim rests on.
The workflow eval itself needs one caveat. Accuracy is not measured against ground truth. It is agreement with the average answer of GPT-6 Astra and Claude Fable 5.1 at high thinking, on four workflows written by TypeSafe’s own team. Every model runs through TypeSafe’s System One adapter, which forces the LLMs to return probabilities in Jev’s format. TypeSafe lists all of this on the evals site.
Taken together: Jev is a fast, cheap, zero-shot classifier at roughly mid-tier LLM accuracy, with a probability attached to every answer that is closer to calibrated than what an LLM gives you. The probability is the new part. The rest of this post explains why it needed a different kind of model, and what has and has not been shown about it.
Structured output is not new
Typed output from a language model has been a standard API feature for over two years. OpenAI shipped JSON mode in late 2023 and schema-enforced Structured Outputs in August 2024. Anthropic added tool use with JSON schemas in 2024 and strict structured outputs in 2025. Both guarantee the response parses against your schema. If your schema says the label must be one of four strings, you get one of those four strings. That is the same guarantee Jev gives.
The cleanest way to see what Jev adds is to ask all three APIs the same question over the same input and compare what comes back. Take a support ticket classifier with four labels: billing, technical, account, other.
With the Claude API, you describe the schema as a tool and force the model to call it:
response = client.messages.create(
model="claude-sonnet-4-5",
max_tokens=100,
tools=[{
"name": "classify",
"input_schema": {
"type": "object",
"properties": {"label": {"enum": ["billing", "technical", "account", "other"]}},
"required": ["label"],
},
}],
tool_choice={"type": "tool", "name": "classify"},
messages=[{"role": "user", "content": ticket}],
)
label = response.content[0].input["label"]With the OpenAI API, you pass a Pydantic model and get a parsed object back:
class Classification(BaseModel):
label: Literal["billing", "technical", "account", "other"]
response = client.chat.completions.parse(
model="gpt-4o",
messages=[{"role": "user", "content": ticket}],
response_format=Classification,
)
label = response.choices[0].message.parsed.labelWith Jev, the label set is a Choice question and the ticket is the state:
response = client.system_one(
state=ticket,
questions={
"label": Choice(
instructions="What kind of support request is this?",
criteria={
"billing": "Charges, invoices, refunds",
"technical": "Bugs, errors, outages",
"account": "Login, password, profile",
"other": "Anything else",
},
),
},
)
answer = response.answers["label"]
label = answer.choice
probs = answer.probabilities # {"billing": 0.81, "technical": 0.12, ...}All three return a valid label. The difference is the last line. Jev returns a full probability distribution over the four labels as a first-class part of the response, plus a confidence number derived from that distribution. Claude and OpenAI return the label. If you want a probability from them, you have to go and get one, and there are two ways to do that. Both have problems.
One more difference is structural. A Jev call takes a state and any number of questions, each answered independently in the same forward pass. Adding a fifth question to the request barely changes latency. With an LLM, five fields in one schema are generated one token after another, and each field’s tokens are conditioned on the fields before it. That is where most of the speed gap comes from, and it is also why TypeSafe tells you to break a judgement into small atomic questions and combine them in code rather than ask one big question.
Two ways to get a probability out of an LLM
I covered this in a pre-Jev post on Linkedin, so here is the short version.
Method 1: ask for it. Add a confidence: float field to the schema. The model writes a number.
class Classification(BaseModel):
label: Literal["billing", "technical", "account", "other"]
confidence: floatThis is a verbalised probability. It is text the model generated, in the same way it generates the label. Nothing in the model checks that number against anything. In practice it clusters at 0.85, 0.9 and 0.95, and it is the same 0.9 whether the ticket was obvious or ambiguous.
Method 2: read the token probabilities. OpenAI’s API returns logprobs, the log probability the model assigned to each token it emitted, and top_logprobs, the alternatives it considered at each position. Find the tokens that make up the label, sum their log probabilities, exponentiate, and you have the probability the model assigned to the label it chose. Look at the top alternatives at the first label token and you can rebuild a distribution across all four labels.
resp = client.chat.completions.parse(
model="gpt-4o", temperature=0,
logprobs=True, top_logprobs=5,
messages=[...], response_format=Classification,
)
toks = label_token_spans(resp.choices[0].logprobs.content) # tokens covering the label value
label_prob = math.exp(sum(t.logprob for t in toks))
dist = class_distribution(toks[0]) # top-5 alternatives at first token, mapped to labels, renormalisedThis is a real quantity from inside the model: the softmax output at the position where the label was written.
What the two methods look like side by side. I ran both on GPT-4o over 40 short statements: 20 rephrased from earnings calls, 20 synthetic product reviews. Each was classified positive or negative, and for each I recorded the confidence the model wrote and the probability implied by the token distribution.
Across all 40, GPT-4o used six verbalised confidence values: 60%, 70%, 80%, 85%, 90% and 95%. It never wrote 99%.
On 37 of the 40, the chosen label had more than 99% of the token probability mass, often above 99.99%.
On 3 of the 40, the two disagreed: the model wrote 70% to 80% confidence while the token distribution was close to a coin flip.
The model has one vocabulary for talking about uncertainty and a different distribution for producing the answer, and they are not connected to each other.
Three things we call confidence
The word is doing three jobs, and the experiment above only makes sense once they are separated.
Verbalised probability: the model writes
0.85. A generated output like any other.Token probability: the softmax mass on the label tokens, exposed through logprobs. A real internal quantity, but the probability the model preferred that token, not the probability the label is correct.
Calibrated decision probability: a number such that, across many decisions given 0.8, about 80% turn out correct. This is the one software needs, and neither of the first two is trained to be it.
Most of the confusion in the Jev discussion comes from using one word for all three. “Jev returns confidence” and “GPT-4o returns confidence” are both true statements about different quantities.
The logprob method has practical problems of its own. Anthropic does not expose logprobs, so it does not work on Claude. The bookkeeping is fiddly: labels that share a prefix tokenise ambiguously, the label has to be the first field in the schema so nothing else conditions it, and top_logprobs caps at 20 alternatives. And the token probability is not calibrated: 37 of 40 above 99% is not a model that is wrong 1% of the time. RLHF pushes probability mass onto the preferred answer, which is the mode-dropping problem TypeSafe describes in its AI primer.
One research result cuts against the simple story that logprobs are good and verbalised numbers are bad. Tian and colleagues (Just Ask for Calibration, EMNLP 2023) found that on question-answering benchmarks, RLHF models including GPT-4 and Claude gave verbalised confidence that was often better calibrated than their token probabilities, because RLHF sharpens the token distribution so much. So the ranking between the two depends on the task. For a fixed-label classifier at temperature 0, the logprob gives more resolution and ranks cases better, which is what I would use for thresholding on OpenAI models. For open-ended answers, the verbalised number can be the less bad of the two. In neither case was the number optimised against outcomes.
A schema constrains the format of the number, not its meaning. That is the gap Jev is aimed at, and it takes a look at where verbalised probability comes from to see why a logprobs endpoint would not have closed it.
Verbalised probability is learned from people, and people are bad at it
A language model’s verbalised confidence is learned the same way everything else it says is learned: from text written by humans. So the question of whether an LLM’s “0.9” means anything reduces to whether the humans who wrote the training data attached consistent numbers to words like “likely”. The intelligence community has been measuring this since 1964, and the answer is no.
Sherman Kent ran the CIA’s Board of National Estimates. In 1951 an estimate said a Soviet attack on Yugoslavia was a “serious possibility”. Kent meant about 65%. When he asked the colleagues who had signed off on the phrase what number they had in mind, the answers ran from 20% to 80%. His 1964 paper Words of Estimative Probability proposed a fixed scale (”probable” means 75%, give or take 12) and the CIA did not adopt it. Analysts felt numbers were too sharp for the evidence and a fixed vocabulary would constrain the prose.
In the 1970s Scott Barclay and colleagues, writing a decision-analysis handbook for the US Department of Defense, put the question to 23 NATO officers. Each was given sentences like “It is highly likely that the Soviets will invade Czechoslovakia” and asked for a percentage. The dot chart of their answers went into Heuer’s Psychology of Intelligence Analysis and from there into every textbook and slide deck on the subject. Edmund Conrow audited it in 2010 and found the raw data had been lost and later redrawings disagreed with the original; Daniel Hails digitised three published versions in 2026 and found dot counts per phrase varying from 16 to 23 where there should always be 23. So the famous chart cannot be reproduced. What can be reproduced is the 2015 replication by the Reddit user Zonination, who asked 46 people the same 17 phrases and published the raw responses. That is the chart below.
The ranking survives: everyone agrees “almost certainly” sits above “probable” which sits above “unlikely”. The number does not. “We believe” has a median of 70% but the middle half of respondents spread from 60% to 80%, and the full range runs from 5% to 100%. “Highly unlikely” has a median of 5% and at least one respondent at 90%. Hails’ 2026 survey of 99 people found the same medians as the 1970s officers, to within 5 points, and the same 10 to 20 point interquartile spread on most phrases. Fifty years, three populations, same result.
The LLM implication is direct. The model has read millions of sentences where “probably” was written by someone who meant anything from 45% to 90%. When it writes confidence: 0.85, it is producing the kind of number a person would write next to that label in that kind of document. It is a plausible number, not a measured one. And the argument does not stop at verbalised confidence. Token log probabilities are at least a real quantity, but RLHF then trains the model toward the answer a human rater prefers, and raters prefer confident answers. Kalai and colleagues at OpenAI made the same point in Why Language Models Hallucinate (2025): evaluation that scores only right or wrong rewards guessing over abstaining, so the training process itself produces overconfidence. Neither the verbalised number nor the logprob was ever optimised to match the frequency of being right.
Why Jev is different: RLCD
TypeSafe calls its training method Reinforcement Learning for Calibrated Decisions. This is what has been published about it, and it is not much.
What TypeSafe has said. The launch post and the AI primer state the objective and the output contract. RLHF optimises for “writeups and chat responses that human raters prefer”. RLCD optimises for “calibrated decisions: answers with epistemically honest probabilities on System One tasks”. Calibrated is defined the standard way: across many predictions, outcomes given probability 0.2 should occur about 20% of the time, and outcomes given 0.8 about 80% of the time. The primer names two failure modes of RLHF it is trying to avoid: overconfidence, and mode dropping, where the model concentrates probability mass on one style of answer and loses the rest of the distribution. TypeSafe also says the model has a new architecture and a parallel sampler, and that it is non-autoregressive: all questions are answered in one forward pass rather than token by token.
What TypeSafe has not said. The reward function, the training data, the base model, the parameter count, and the recipe. There is no paper. Asked directly on the Latent Space podcast on 22 September 2026 whether one had been published, Almeida answered “No, not yet.” In the same interview he described RLCD as a task rather than an algorithm: RLHF now means the task of instruction following regardless of whether a system uses the PPO recipe from the original paper, and RLCD is meant the same way, as a new training target. On what happens when calibration is wrong: “We have a report issues button. Complain to us in Discord.” The acronym also collides with an unrelated 2023 method, Reinforcement Learning from Contrastive Distillation, which is why an arXiv search for RLCD returns the wrong paper. Every third-party explainer on RLCD is paraphrasing the same three sentences from the docs. TypeSafe does not claim RLCD improves accuracy, only calibration, and the eval numbers agree: Jev’s accuracy sits with mid-tier LLMs.
What can be inferred. The mechanism is not a mystery even if the recipe is. If you want a model’s stated probability to match its hit rate, you score it with a strictly proper scoring rule, a loss function whose minimum is reached only when the reported probability equals the true probability. Brier score and log loss are the two standard ones. Train a policy with reinforcement learning where the reward is the negative Brier score of its stated probability against the verified outcome, and the model is penalised for saying 0.95 on things it gets right 70% of the time, and equally penalised for saying 0.6 on things it gets right 95% of the time. There is no reward for sounding sure. This has been done in the LLM setting at least twice in the last 18 months. Stangel and colleagues’ Rewarding Doubt (ICLR 2026) uses a proper scoring rule as the RL reward so the model’s stated confidence tracks factual correctness. Damani and colleagues’ RLCR (2025) adds a Brier-score term to the correctness reward of a reasoning model and shows calibration improving without accuracy falling. The open-source Laya model, built as a Jev alternative, documents training against the Brier score explicitly. It would be surprising if RLCD were doing something fundamentally different in its objective, whatever the architecture underneath.
What that objective does to the two kinds of probability from the earlier section is the point. In a chat model, the verbalised probability (text) and the token probability (softmax) are separate quantities and neither is trained to match outcomes. In Jev there is no verbalised probability, because there is no text, and no token probability in the chat-model sense, because there is no token-by-token generation. The model’s only output is a distribution over the permitted answers, and that distribution is what the reward is computed on. My first framing of this was that RLCD aligns the model’s internal probability with its verbalised one. That is not right: Jev removes the split rather than reconciling it. The probability is not commentary about the answer, it is the answer. That is the difference between a model that has read a million documents where people wrote “likely” and a model that was scored on whether 70% of its 0.7s came true.
Two caveats keep this honest. First, calibration is a property of a distribution of inputs. A model calibrated on TypeSafe’s training data is not automatically calibrated on yours, and the independent out-of-distribution test showed exactly that: ECE of 0.02 to 0.03 on public benchmarks, 0.107 on synthetic tickets with a rule the text could not reveal, and the direction of the error flipping between question types. Second, a calibrated model can still be wrong, and can be confidently wrong on individual cases. What calibration buys you is that the confidence number is a usable signal for routing, not that the answer is right. The practical fix reported across several independent tests is cheap: 50 to 300 of your own labelled cases and one fitted temperature parameter reduced ECE by up to 74%. That is the same post-hoc calibration you would apply to any classifier, and it works on Jev because there is a real distribution to recalibrate.
Where it fits
The use case is not “replace your LLM”. It is the decision points inside a workflow where you currently either hard-code a rule that is too brittle or call an LLM that is too slow and too expensive for a yes/no. Route this ticket. Is this alert worth an analyst’s time. Does this invoice match the purchase order. Should this agent step be checked by a person before it runs. Each of these is a Choice, a Score or a Noul, and each needs a probability so the code can decide when to act and when to escalate.
A calibrated probability turns that into a threshold problem. TypeSafe’s confidence docs give the pattern: act automatically above one threshold, confirm with a person in the middle band, refuse to act below a floor, and set the thresholds by the cost of being wrong rather than one number for everything. A read-only action can run at 0.6 confidence. Approving a transfer needs 0.9 and a confirmation step. The point of calibration is that the threshold means the same thing next week as it does today, provided you pin the model version. With an uncalibrated model the threshold is a guess that drifts.
The economics follow from the same number. Suppose a workflow makes 100,000 decisions. If 95% clear the threshold and run automatically, 4% escalate to a reasoning model and 1% go to a person, the cost of the workflow is set almost entirely by those last 5%. The model’s uncertainty is deciding how much expensive intelligence the system buys. Over time that matters more than the per-token price, and it only works if the 95% that were waved through are as reliable as the number said they were. That is the operational meaning of calibration.
The independent results so far say where this works and where it does not.
On cost, the right comparison depends on what you would otherwise use. Against a frontier reasoning model on a multi-question workflow, two orders of magnitude is real. Against a small fine-tuned BERT or a gradient-boosted model you already have, Jev is neither faster nor cheaper, and it is unlikely to be more accurate on that model’s own distribution. What it saves is the training set. Jev is zero-shot from a label description, so it fits the cases where you have a new decision to automate and a few hundred labels at most, not the cases where you have a hundred thousand.
None of the individual pieces is new. Zero-shot classification, structured output, proper scoring rules and reinforcement learning for calibration all existed before Jev. What is new is a model whose whole job is the decision and the probability, with no text in between. Whether TypeSafe’s recipe holds up is open until it is published or reproduced, and the calibration claim needs testing on more distributions than four internal workflows and a handful of public benchmarks. The requirement it addresses is not in doubt: a model that makes unattended decisions in software has to report how often it is wrong, and that number has to come from being scored against outcomes, not from reading how people describe uncertainty.
The experiment to run
The comparison in section two is qualitative. The quantitative version is the same task and the same output contract across all three APIs, over a labelled dataset, scored on calibration and not only accuracy.
Use one label set and one prompt. For OpenAI and Claude, enforce
{label, confidence}with structured outputs. For OpenAI, also capture the label logprobs. For Jev, express the same label set as a Choice question and keep the returned distribution.Run all three over the same labelled set. The 40 sentiment statements from the earlier experiment are a start, but a few hundred cases per label are needed for the calibration buckets to mean anything.
Score each signal (verbalised, logprob, Jev) on accuracy, Brier score, log loss and expected calibration error, and draw the reliability curve: within each confidence bucket (50 to 60%, 60 to 70%, and so on up to 90 to 100%), what fraction was correct.
Perturb the inputs and rerun: rephrase without changing meaning, add irrelevant IDs and metadata, reorder fields, move from synthetic reviews to earnings-call language. A calibrated model should keep its reliability curve under these changes; a sharpened one should not.
Record latency and cost per case from the same machine and the same time window.
The methodological wrinkle is that the three calls are not the same operation. OpenAI and Claude generate a JSON document that happens to contain a number; Jev returns a distribution from a narrower interface. That difference is what is being tested. The expected result, based on the independent tests so far, is that OpenAI and Claude match or beat Jev on accuracy, Jev wins on latency and cost by one to two orders of magnitude depending on the baseline, and Jev’s probabilities sit closer to the diagonal on the reliability curve, with the gap narrowing after one temperature fit on the LLM logprobs. I have not filled that table with guessed numbers. It needs to be run.
What this means for business decision-making
Most of what is written about Jev is about the model. The useful conclusions are about how decisions get automated, and they apply whether or not you ever call TypeSafe’s API.
Check what your confidence numbers are. If an LLM workflow you run today routes on a
confidencefield the model wrote, it is routing on a generated number that clusters at 0.85 to 0.95 regardless of the case. Pull 200 cases, compare the stated confidence with the outcome, and look at the result before trusting the threshold.Benchmark the probability, not only the accuracy. The question for any decision model is: when it says 60%, 70%, 80% and 90% on my data, how often is it right? A reliability curve on your own labelled cases answers that. Accuracy alone does not tell you whether the model can be left unattended.
Set thresholds by the cost of being wrong. One threshold for a workflow is wrong. A decision that is cheap to reverse (show a screen, tag a record) can run at lower confidence than one that is not (approve a payment, close an incident). The threshold is where the business encodes its risk tolerance, and it belongs in code where it can be reviewed and changed.
Break judgements into small questions and keep the rules in code. “Should this invoice be paid?” becomes: does the amount match the order, was the delivery confirmed, is the vendor on file, each answered separately, with the approval logic written as ordinary rules. This is the pattern TypeSafe’s evals are built on, and every model in those evals scored better with it than with one large prompt. It also makes the workflow auditable: a changed policy is a changed line of code, not a rewritten prompt.
Budget for the escalation tail. In a cascade, the cost of the workflow is set by the share of decisions that fall below the threshold and go to a reasoning model or a person. Track that share. A calibrated decision layer lets you predict it; an uncalibrated one hides it until the human review queue fills up.
Recalibrate on your own data, whatever the vendor. Fifty to three hundred labelled cases and one fitted temperature parameter cut Jev’s calibration error by up to 74% in independent tests. The same fix applies to any model that exposes a real distribution. It is cheap and it should be a standard step before any threshold goes live.
Know where a zero-shot decision model fails. A rule that is not in the text (internal policy, tacit knowledge, an exception list) will be answered confidently and wrongly. Put the rule in the state or in the code. Many-class problems, non-English text and adversarial content are documented weak spots for Jev specifically.
Treat the model version as part of the decision. Thresholds are fitted to one model. TypeSafe has said it will ship new models quickly and is not promising long-term support. Pin the version, and re-run the reliability curve when it changes. This applies equally to LLM-based classifiers, where a silent model update moves the logprob distribution.
The short version: the value of a decision model in a workflow is the probability it attaches to each decision, and that probability is only worth what it has been tested to be worth on your data. Jev is the first vendor to make that number the product. The discipline of measuring it is what businesses should take from the launch, whichever model they end up using.
Sources
TypeSafe AI, Introducing System One Models & Jev, 15 September 2026
TypeSafe AI, Workflow evals, AI primer, Confidence, Introduction
LiteLLM, JEV Classifier: 5.43x as Fast as Haiku, 96% Lower Cost
scienthoon, Independent calibration test of Jev, 4,621 calls
xbill, Jev After Eight Days of Independent Tests, review of 14 preprints, 104 repositories and 33 blog posts
Delip Rao, JEV vs. LLMs as Rubric Judges, arXiv 2609.29769
Calibrated Decision Models for Autonomous Penetration-Testing Harnesses, arXiv 2609.28940 (Laya’s Brier-score training)
Damani et al., Beyond Binary Rewards: Training LMs to Reason About Their Uncertainty, arXiv 2507.16806, 2025
Kalai et al., Why Language Models Hallucinate, arXiv 2509.04664, 2025
Sherman Kent, Words of Estimative Probability, Studies in Intelligence, 1964
Daniel Hails, The CIA was “Probably” Right, April 2026 (Barclay 1977 digitisation, Conrow 2010 audit, 2026 survey n=99)
Zonination, Perceptions of Probability, 2015 survey raw data, n=46
DataCamp, Jev: TypeSafe’s System One Model
systemonemodels.org, RLCD explained (arXiv search result for the term)
Tian et al., Just Ask for Calibration, EMNLP 2023
Stangel et al., Rewarding Doubt: A Reinforcement Learning Approach to Calibrated Confidence Expression of Large Language Models, ICLR 2026
Latent Space podcast, interview with Diogo Almeida, 22 September 2026 (transcript summary)





