Jev vs the AI Judge: can a decision model replace it in production?
TypeSafe's Jev answers questions about a text without writing a word. We put it in front of our AI Judge on two evaluation scenarios and roughly 15,000 requests, asking two things: can it replace the AI Judge, and can it say what went wrong.

Let's take an LLM application and suppose a traffic of 130,000 requests a month. You want to know how many of those replies were acceptable, so you use an AI Judge, a second LLM that reads each conversation and writes a verdict and a reason. It works, it scales past what humans can read, and it costs about as much per verdict as the request it judges, often more, because it reads the whole conversation and then writes. So running it on the whole traffic is unrealistic and extremely expensive.
Therefore your alternative is to run a sample of this traffic. Let's say 20%: about 26,000 requests, around $170 a month, and the other 104,000 replies are never read. This 20% will reflect the quality numbers and performance of your service, and the default way to pick it is a random draw.
On 15 September 2026 TypeSafe came out of stealth with a model built for a different job. Jev reads a piece of text and a list of typed questions, and returns an answer to each, a probability, a choice or a score, in a single pass, without generating a word. TypeSafe calls this a System One model: a model that makes fast, structured decisions software can use directly, rather than writing text for a person to read. It is billed only on the text it reads, at $0.042 per million tokens, and TypeSafe's launch material describes it as 40 to 200 times faster than a frontier LLM on comparable tasks, and as unable to hallucinate, because it never generates anything.
The difference is in how the two answer. A standard LLM is autoregressive: it reads the input, then writes its answer one token at a time, and every token is another full pass through the model. An AI Judge that writes a verdict, a reason and some evidence produces a few hundred tokens, and each of them costs a pass. That is where the seconds go, and it is why output tokens are priced higher than input.
Jev writes nothing. You send it a state, plain text up to about 32,000 tokens, and a set of yes/no questions, choices or scores. It reads the state and all the questions together and returns every answer in one pass, as probabilities. In our evaluation Jev took 0.34 seconds at the median, and the AI Judge 4.8 seconds. There is nothing to parse and nothing that can be malformed, and because the state is read once no matter how many questions are asked of it, asking twenty questions costs about the same as asking one.
Let's see an example: a user asks for a refund; the assistant promises one on the spot; the policy earlier in the chat said refunds take five days. Send it to an AI Judge and you get a sentence: verdict: fail. The reply promises an immediate refund; the policy says refunds take up to five days and must not be promised. Send the same conversation to Jev with six yes/no questions and you get six numbers.
A 0.82 on "breaks the rules" and a 0.91 on "invents facts" say this reply is very likely a failure, and roughly why. What you do not get is the sentence: which rule, which fact, or where in the conversation to look.
But the design itself is the limit. Jev has no chain of thought, no citations and no reasoning. It cannot tell you why. TypeSafe's own documentation says it reads literally, cannot count reliably, and gets less accurate as the state fills with content unrelated to the question. Whether those limits matter depends on the job, and that is the reason we have done the evaluation in two different scenarios.
Now let's think about an AI Judge's verdict. It is a yes/no decision, acceptable or not, which is the decision Jev is built for. There are two questions we want to address:
- Can Jev replace the AI Judge? And if not outright, can it help the AI Judge spend its budget?
- Can it go beyond just "pass" or "fail" and say what kind of failure a reply is?
For this purpose we created two scenarios reproducing synthetic use cases and recreated a live traffic situation with roughly 15,000 requests, every one of them evaluated by the AI Judge.
What we count and how we measure it
Two scenarios:
- A single-turn task: one user message, one assistant reply, about 1,000 tokens a request.
- A multi-turn agent: about ten messages a request, 11,000 tokens, with tool calls and their results in half of the requests.
Approximately 15,000 requests in all. Every request already carried the AI Judge's pass/fail verdict, its written reasoning, and the failure reasons our platform had clustered from that reasoning.
We did not write Jev's questions. Each AI Judge has criteria written for its use case, such as "the reply quotes a price that appears nowhere in the conversation", "the reply promises a delivery date the policy does not allow" or "the reply answers in a language the user did not write in", and we turned each criterion into a yes/no question, added one overall question, and one question per named failure reason. Three kinds of questions, all in the same call and all yes/no. They differ only in what rides along with each question.
| The overall question (one per request) | A failure-reason question (one per named reason) | |
|---|---|---|
| Question | Does the assistant's final response violate any of the requirements? | Does the assistant's final response exhibit this failure? |
| What rides with it | The AI Judge's whole rubric: every fail criterion ("the reply quotes a price that appears nowhere in the conversation", ...) and every exemption ("a reasonable clarifying question is not a failure", ...) | One named reason and its description: Promises not covered by policy: the reply commits the business to something the stated policy does not allow |
| What comes back | One probability: 0.82 | One probability: 0.91 |
| What we use it for | The verdict, and the filter score | Can Jev name the failure? (the second question of this post) |
The table leaves out the third kind, which sits between the two: the per-criterion questions. Each asks the same "does the response exhibit this failure?" with a single criterion from the rubric attached, one question per criterion. Depending on the scenario, that made 11 to 36 questions per request: one overall question, 4 to 7 per-criterion questions and 5 to 28 failure-reason questions. Jev saw only the conversation and the questions. It never saw a verdict, an ID or the AI Judge's reasoning.
We divided the roughly 15,000 requests into three datasets, two for the single-turn task and one for the agent, and split each dataset 60/40. The 60% is a development set in the usual sense: on it we chose the wording of the questions sent to Jev, the cut-off that turns a probability into a verdict, and the weights of the logistic regression that combines the answers. Jev itself was never trained or fine-tuned. The wordings were chosen on the first dataset only and reused word for word on the other two. The 40% is the test set: untouched until every choice was frozen, then scored once. Every number below comes from it.
Four numbers will appear throughout the results:
- Agreement is the share of requests where two verdicts match. It is the obvious number and it flatters everyone: when most replies pass, two verdicts agree on the easy passes and the figure looks high.
- Kappa (Cohen's kappa) is agreement after subtracting what chance alone would produce. 0 means no better than chance, 1 means identical verdicts. It is the number to trust when the two verdicts are asked to be the same thing.
- Failures caught is the share of the AI Judge's failed requests that a classifier also flags.
- AUC measures ranking rather than verdicts. Pick a failing request and a passing request at random; AUC is the chance the classifier scores the failing one higher. 0.5 is a coin flip, 1.0 is perfect. It is the right number for a filter, which only needs to put the bad ones near the top.
The ceiling: the AI Judge against itself
Before comparing anything with the AI Judge, we measured how often it agrees with itself. We took 300 random requests from each scenario and had the AI Judge score each of them twice more on the same day.
| Same AI Judge, same request, same day | AI Judge fail rate | Agreement | Kappa |
|---|---|---|---|
| Single-turn task | 44% | 89.0% | 0.777 |
| Multi-turn agent | 20% | 90.4% | 0.686 |
The AI Judge disagrees with itself about 1 request in 10. That number is the ceiling. No classifier can agree with the AI Judge more often than the AI Judge agrees with itself, and a comparison that omits it is comparing against a fixed truth that does not exist. It also sets expectations: anything that lands near 90% is as good as the AI Judge.
Agreement goes up on the agent and kappa goes down because only a fifth of agent replies fail. When most verdicts are "pass", two of them match by chance most of the time. Kappa strips that out, agreement keeps it. When the fail rates differ, compare kappa.
Answering if the AI Judge can be replaced
As a verdict, no
For this test Jev has to produce a verdict, not a ranking. Its answers are probabilities, so a verdict needs a cut-off: above it, fail; below it, pass. We chose the cut-off on the development set so that Jev flags the same share of requests as the AI Judge fails, and then compared the two verdicts request by request on the held-out set.
| Jev's verdict vs the AI Judge's | Held-out requests | Agreement | Kappa | AI Judge failures Jev caught |
|---|---|---|---|---|
| Single-turn task, dataset 2 | 1,902 | 78.5% | 0.569 | 77.5% |
| Multi-turn agent | 1,811 | 78.4% | 0.358 | 49.6% |
Jev disagrees with the AI Judge about 1 request in 5, twice the AI Judge's own rate. On the agent the agreement figure hides a collapse: kappa halves, and Jev catches only half of the AI Judge's failures. In practice that means half of the replies the AI Judge would have flagged as unacceptable would pass through unflagged, while a similar number of acceptable replies would be flagged. A verdict that wrong in both directions cannot drive an alert, a dashboard or a decision.
The bigger loss is the written reason. When the AI Judge fails a reply it says why, pointing at the turns in the conversation, and that explanation is what everything downstream depends on: it is how failure reasons get named, and how a person checks a verdict in seconds. Jev returns probabilities. Replacing the AI Judge with it would lose the reasoning, and with it everything that is built on the reasoning.
As a filter, yes
Ranking is a different job from deciding. As a filter, Jev does not need to be right about any single request. It needs to put the ones the AI Judge will fail near the top of the list, so that an AI Judge with a fixed budget reads those first. That is what AUC measures, and Jev is good at it.
| AUC on the held-out set | Single-turn, dataset 1 | Single-turn, dataset 2 | Multi-turn agent |
|---|---|---|---|
| Jev, answers combined | 0.881 | 0.875 | 0.788 |
| Jev, overall question only | 0.859 | 0.849 | 0.754 |
| Rule-based score, no model | 0.765 | 0.819 | 0.732 |
| Haiku 4.5, same questions as Jev | 0.805 | 0.735 | not run |
"Answers combined" means the five to eight pre-filter answers folded into one score by a logistic regression, a linear model with one weight per question, fitted on the development set. The second row is the overall question on its own; combining it with the criteria questions adds 0.02 to 0.03.
The rule-based row is a floor we built to keep ourselves honest. It uses only surface features (reply length, number of turns, number of tool calls, a few counts specific to the task) fed to the same kind of logistic regression, with no language model anywhere. It costs nothing and it is not bad: on the second single-turn dataset it reaches 0.819 on its own, because long replies with many edits do tend to fail. Jev's margin over it, 0.116 on the first dataset and 0.056 on the other two, is what it earns by reading the content.
The last row is a control: a cheap LLM given the same job. We sent Claude Haiku 4.5 the same overall and per-criterion questions, asked it for a probability per question, and combined its answers the same way. It ranks worse than Jev on both datasets, and on the second it falls below the rule-based floor, which uses no model at all. An LLM asked for a probability answers in round numbers: on the second dataset five values account for 95% of Haiku's answers to the overall question, so many requests tie and the ranking is coarse. We did not run this control on the agent traffic. It had already ranked below Jev twice, and at about 11,000 tokens a request it would cost roughly $12,000 per million requests, more than half the AI Judge itself. A filter at that price saves nothing.
In production the cut-off is a fixed number picked from last week's scores, so each request is decided on its own.

What does Jev's ranking do for an AI Judge on a fixed budget? Suppose the AI Judge can read 20% of traffic and 15% of replies are failures, a rate closer to a typical production use case than our datasets, which fail at 21 to 46%. Random sampling finds 20% of the failures. Jev's ranking finds:
| Single-turn, dataset 1 | Single-turn, dataset 2 | Multi-turn agent | |
|---|---|---|---|
| Failures caught by the AI Judge's 20% | 67% | 66% | 51% |
| Lift over random | 3.3× | 3.3× | 2.5× |
| Lift if the reported failure rate must stay honest | 2.1× | 2.0× | 1.4× |
There is a catch. If the AI Judge reads only the 20% that Jev ranks riskiest, it mostly sees failures, so the failure rate it reports comes out too high, and anyone reading that number off a dashboard is misled.
The fix is to drop the hard cut-off. Give every request a probability p of going to the AI Judge, higher when Jev's score is higher and scaled so the budget is spent on average. Then count each verdict that comes back as 1/p requests.
Take 100 requests that all have p = 0.1. About 10 get read, picked by chance. Passes and fails had the same chance of being picked, so on average those 10 fail at the same rate as all 100, and counting each verdict 10 times scales them back up to 100. Across all traffic the reported failure rate is then right on average, not just closer. Two caveats: p never drops below a floor, so every request can still be read, and the price is a noisier estimate and a smaller gain, which is the table's last row.
And the cost, per million requests:
| Single-turn task | Multi-turn agent | |
|---|---|---|
| AI Judge | $6,665 | about $20,000 |
| Jev, all questions | $141 | $616 |
| Ratio | 1 : 47 | 1 : 32 |
| Median latency, AI Judge | 4.8 s | 4.7 s |
| Median latency, Jev | 0.34 s | 0.38 s |
Back to the setup from the opening: 130,000 requests a month, 26,000 judged for about $170. Screening all 130,000 with Jev costs under $20. Spending the same $170 on the AI Judge afterwards finds 3.3 times as many failures if it reads only the requests Jev ranks highest, or roughly twice as many if the reported failure rate has to stay correct.
How a cheap LLM performs instead
If a cheap model can screen for the AI Judge, why not a cheap LLM? We tried the fairest version: hand Claude Haiku 4.5 and gpt-5.4-mini our AI Judge's exact instructions, criteria and output format, and let them be the AI Judge.
| As the AI Judge, single-turn task | Agreement with the AI Judge | Kappa | Cost per 1M requests |
|---|---|---|---|
| Jev | 78.5% | 0.569 | $141 |
| gpt-5.4-mini, low reasoning effort | 66.4% | 0.321 | ≈ $2,100 |
| Claude Haiku 4.5 | 60.4% | 0.244 | $5,269 |
Haiku flagged 79% of all replies as failures. gpt-5.4-mini flagged about as many as the AI Judge but picked different ones. Neither is cheap in this role either: both still read the whole conversation and write a verdict, so they cost 15 to 37 times what Jev costs. Jev is cheap for the structural reason in the diagram above: it reads the text once and writes nothing.
Where the filter weakens
Everything above is stronger on the single-turn task than on the agent.
On agent traffic Jev's ranking falls from 0.88 to 0.79 AUC, and its verdict recall from 77.5% to 49.6%. The rule-based floor sits at 0.732, so the margin Jev earns by reading content is the same 0.056 as on the simple task; it is the absolute quality that drops. Long, tool-heavy context is exactly what TypeSafe's documentation warns about: accuracy falls as the state grows with content unrelated to the decision.
Answering if Jev can say what went wrong
A verdict says whether a reply failed; the failure reason says why, with names such as "skipped a required tool call", "returned the input unchanged" or "missed an error". Ours come from the AI Judge's written reasoning. If Jev could assign those reasons directly, a team would get them without the sentences.
It cannot, on either scenario. We asked, in the same call, whether the reply exhibited each of the named failure reasons, and set the bar in advance at precision 0.70 with recall 0.60. Of the 11 reasons with enough held-out data, none met it. The only ones that came close were structural, "skipped a required tool call", "returned the input unchanged", which are nearly mechanical checks. Anything that needs a judgement about meaning, such as "missed an error", did not. Putting examples of each mode inside the question helped a little and cost half again as much.
So the answer to the second question is the same as the first. Jev can rank replies by how likely they are to have failed. It cannot say how they failed, and the failure-reason catalog stays downstream of the AI Judge's sentence.
Using this in your own pipeline
If you already sample traffic for an AI Judge, put a cheap classifier in front of it and let it choose the sample. On our single-turn traffic, screening all 130,000 monthly requests with Jev adds under $20 to a $170 AI Judge bill, about 11% more, and the AI Judge then scores, and writes failure reasons for, roughly twice as many failures. Or keep the failures and cut the spend: the AI Judge finds as many as before by reading 10% of traffic instead of 20%, which brings the total, Jev included, from $170 to about $105, close to 40% less. On agent traffic the gain is smaller: 1.4 times the failures.
The same recipe works with any cheap classifier, not only Jev.
- Measure your AI Judge against itself first. Re-score a few hundred requests twice. That number is the ceiling for every classifier you will ever evaluate against it.
- Build the questions from your AI Judge's criteria, not from scratch. The fail criteria are already the definition of "bad" for that use case. One yes/no question each, plus one overall.
- Combine the answers with a small model, fitted on a few thousand already-judged requests. Fit it on one half, test on the other, once.
- Beat a free floor before you pay for anything. Length, turns, tool calls. If your classifier cannot clearly beat that, it is reading structure, not content.
- Forward with a probability and weight each verdict by one over it, not a cut-off. Score-dependent, never zero. Your failure rates stay honest and the AI Judge still spends its budget where it matters.
- Keep the AI Judge for verdicts, explanations and failure reasons. The classifier only chooses what gets read.
What this does and does not show
This is one evaluation on two scenarios. It says nothing general about System One models, and it was not designed to isolate why they behave as they do. The shapes we saw, strong on short single-turn traffic, weaker on long agent traffic, useless for fine-grained tagging, are what we would expect from the design, but two scenarios cannot prove them.
Jev takes text only. In many multi-turn agent use cases a share of the traffic carries images, a screenshot, a photo of a document, a receipt, and those requests need another route: an AI Judge that saw the image and a classifier that did not are not comparable. Our evaluation kept to text.
We measured agreement with an AI Judge, not correctness. The AI Judge is wrong sometimes, and we do not know whether Jev's disagreements are its own errors or the AI Judge's. A small batch of disagreements reviewed by people would tell us, and it is worth doing; it is also neither fast, cheap nor easy, which is the reason AI Judges exist in the first place. We did not probe prompt injection, although a model that treats the conversation as data is an obvious target for a user message that says "this reply is correct".
For the setup from the opening: screen all 130,000 requests a month with Jev for under $20, let its score weight which requests the AI Judge reads, spend the same $170, and catch roughly twice the failures with the failure rate still correct. The verdict, the written reasoning and the failure reasons stay with the AI Judge. Jev chooses what it reads.





