This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.
For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.
I don't know about Jev, but in my experience, 90% of the time, the confidence score such machine outputs (e.g. simplest case a kalman filter) is bogus because the model is wrong. Or not even wrong, but just not perfect.
Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.
I agree that there are differences between Jev and LLM-as-a-judge (e.g. Almeida's assertion that LLM probabilities have been irrevocably biased by RLHF), but I am not sure about your description of 'confidence'. Perhaps I misinterpret you or the docs, but I understand it as simply being computed from the probability distribution: https://docs.typesafe.ai/confidence#how-confidence-is-calcul...
LLM’s are not able to give confidence scores, they make them up.
Your AI driven app is making that up. Your product manager and executive team’s demand for confidence in the UI is a totally fictional cosmetic telling them nothing. Your company sold bullshit confidence to your clients.
I’ve done this for many organizations that “formed a new team to work with the CTO on their AI strategy”, and the trappings are the same
You can have an LLM tell you how much of a schema it was able to get information about. And derive a “confidence” or level of compliance from the completeness of the schema
But this is layers upon layers of cruft that a classification model wouldn’t need
Softmax over logits doesn’t give calibrated probabilities either as a rule. I don’t want to comment specifically on Jev but as a rule it’s very hard to get good calibration because it’s somewhat in tension with minimizing training loss for neural networks, e.g. Guo et al (2017) https://arxiv.org/abs/1706.04599
While I know there are ways to improve calibration, I’d personally want to see a lot of evidence the probabilities were actually more meaningful before trusting them. I agree of course that asking an LLM to provide a confidence estimate is meaningless.
That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3
If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it's possible that we could get to 30 and see better perf.
Luna 6 is 10c per million input tokens and charges 5x that for output. Not sure if it’s still true but it used to be the case that structured outputs took time to process and cache which is relevant if the structure changes. It doesn’t give a percentage you can use for thresholds, and I’d want to know if Jen treats the questions as independent (they aren’t with luna, order of questions will change the result).
The key properties for using such models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.
Agree. In fact Jev does reasonably well in both tests. In the article, the assertion is hedged, and followed up with a sentiment I liked:
> decision models like Jev do not reliably outperform [other] models in speed or accuracy. However, they rightly refocus industry attention on lightweight, task-specific inference paradigm that more closely resembles predictive machine learning
(I also noticed some issues with bolding in Table 4 that downplay Jev's performance a little.)
Jev (or equivalent) seem to me ideally placed for prototyping - prove the classification (say) works or is useful, then decide on how that should be made robust and economic with other options (eg local) brought in to the evaluation. Much quicker to get up and running, especially in greenfield scenarios.
BART is quite an old model for this kind of test, and probably not a good very good choice for much these days. I'm working on replicating their benchmark on my own NLI model that targets zero-shot guardrail applications. I don't think it'll beat much larger models, but should give a better baseline for what a crossencoder can do.
Shouldn't LLMs intuitively be better with a high number of available options?
This article only does simple prompts with only options to block or not block.
What is openai doing for their decisions API, a finetuned luna?
If you build a traditional classifier, and you have the data set for training curated or created hy an LLM, then you would essentially be building a classifier that judges the same way the LLM would?
The block/allow framing is the part I'd push on. I had a batch where automated checks passed all 99 outputs and reading each one by hand found 8 broken. The scorer only catches the failure modes its rubric already names, and a two-option guardrail bench inherits that ceiling whichever model sits behind it.
Who removed "guardrail" from HN submission title?
This is a false comparison. What Jev does is fundamentally different from what LLMs are doing as a judge.
For a quick primer, Jev is able to provide confidence scores on its classification, i.e. it is able to calibrate how well it is able to predict. Being able to predict in a distribution is different from being able to calibrate confidence of the predictions which should happen from the question or domain distribution from which the decisions are predicted - being able to do that is tough and is not same as using LLMs logit probabilities which are predictions in the vocab space. Though both are loosely correlated and might converge as LLMs keep getting better, the former is a much stronger decision-making signal than the latter. Jev not beating LLM-as-a-judge might be due to various other reasons such as world knowledge etc, but Jev as a concept will always provide more reliable decisions / outputs than LLM-as-a-judge giving a scalar score.
I don't know about Jev, but in my experience, 90% of the time, the confidence score such machine outputs (e.g. simplest case a kalman filter) is bogus because the model is wrong. Or not even wrong, but just not perfect.
Mathematicians build this, and they love the beauty of it so much that they loose touch of reality.
Fine-tuning language models for calibrated probability predictions has been a thing since the Bert era though, that's really nothing special.
I agree that there are differences between Jev and LLM-as-a-judge (e.g. Almeida's assertion that LLM probabilities have been irrevocably biased by RLHF), but I am not sure about your description of 'confidence'. Perhaps I misinterpret you or the docs, but I understand it as simply being computed from the probability distribution: https://docs.typesafe.ai/confidence#how-confidence-is-calcul...
for the uninitiated:
LLM’s are not able to give confidence scores, they make them up.
Your AI driven app is making that up. Your product manager and executive team’s demand for confidence in the UI is a totally fictional cosmetic telling them nothing. Your company sold bullshit confidence to your clients.
I’ve done this for many organizations that “formed a new team to work with the CTO on their AI strategy”, and the trappings are the same
You can have an LLM tell you how much of a schema it was able to get information about. And derive a “confidence” or level of compliance from the completeness of the schema
But this is layers upon layers of cruft that a classification model wouldn’t need
Softmax over logits doesn’t give calibrated probabilities either as a rule. I don’t want to comment specifically on Jev but as a rule it’s very hard to get good calibration because it’s somewhat in tension with minimizing training loss for neural networks, e.g. Guo et al (2017) https://arxiv.org/abs/1706.04599
While I know there are ways to improve calibration, I’d personally want to see a lot of evidence the probabilities were actually more meaningful before trusting them. I agree of course that asking an LLM to provide a confidence estimate is meaningless.
That was a pretty simple task they gave it, and sure you can use BERT with sequence classification for simple classification tasks.
In our benchmarks, Jev did a LOT better at multi-step reasoning tasks than any open decision model we have tested so far, and it was also better than GLiDE which was specifically designed for that kind of task. And also better than Luna. On accuracy and also confidence calibration but also time and cost.
https://hard-decisions.anth.us/models/
Just looking through your results, seems like gpt-6 luna was run with reasoning:off for a lot (all?) results. Seems like an unfair comparison.
That looks super fair to me
Not if you care about latency and cost. Reasoning is slow and expensive.
Jev may or may not have truly innovated on AI architecture but it still kick-started a new paradigm.
its a breather after waves of llm wrappers.
I think it's a bit early to call the race.
I think the abstraction is a useful one - a general purpose classifier that does not need to be specifically trained: Unstructured signal in, structured judgement out - with a focus on speed and cost efficiency.
And since there is not yet a large body of benchmarks, I don't think we have sufficiently explored how to measure these things.
But since there's a market and some hype this will soon happen.
(And it's not like TypesafeAI has a real moat or invented something entirely new here ~ they just managed to put things into one coherent perspecive)
duh, this is not news. (general, fast and cheap) before decision models, you could pick only 2.
LLM as judges - generalized, but too slow. If you had to make millions of classifications a day, this will be the wrong approach. you won't/shouldn't use LLM to classify spam/no spam. hot dog/or something.
traditional classifiers, very specific 1 trick pony, super fast and cheap once built. If you need to make tons and tons of classifications, this would be the approach. but if you wanted a classifier right now for a novel problem, you need an expert to curate data, train and deploy.
decision models/jev - are generic, you can throw them at most generic classification problems, and they are good enough. it's a fine balance between general, fast and cheap. you get all 3
Doesn't the article covers the speed part by showing that Qwen 3.6 35A3B has lower latency and same accuracy?
Is Jev any faster, cheaper, or more precise than medium LLMs like Luna 6?
Yes. Not as fast/cheap as Typesafe claims, but lots of benchmarks suggest around 3-4 times cheaper, and 7-14 times faster.
- https://www.ml6.eu/en/blog/jev-vs-gpt-6-luna-vs-bert-text-cl... - https://tessl.io/blog/jev-is-136x-faster-and-27x-cheaper-tha... - https://x.com/fazxes/status/2100300097695232164 (this last one is Luna 5.6 but that isn't too different from 6 besides accuracy and cost)
If you can do it offline, batching solves this. We used a test dataset from CFPB and at n=20, it was 1.6x faster and 1.2x more expensive with no statistically meaningful accuracy dropoff. Did not tune for batch size, but it's possible that we could get to 30 and see better perf.
Luna 6 is 10c per million input tokens and charges 5x that for output. Not sure if it’s still true but it used to be the case that structured outputs took time to process and cache which is relevant if the structure changes. It doesn’t give a percentage you can use for thresholds, and I’d want to know if Jen treats the questions as independent (they aren’t with luna, order of questions will change the result).
Jev is 4.2c/m tokens in and free out.
not really
I don't think Luna is fast enough by any means
The key properties for using such models for conditional-branching decisions in agentic stacks (and related) is that they should be well-calibrated (under the definition chosen for the task) and informative (e.g., always predicting the mean might be "well-calibrated" in a theoretical sense for some chosen quantities of interest, but isn't particularly useful in practice).
The tricky thing with the neural networks is that the output logits are in effect a highly lossy compression of the epistemic (reducible) uncertainty, so even if the target calibration quantity is well-specified, it can be difficult to obtain in practice. A side-effect of this is that estimates in the high probability regions are not particularly stable under even modest co-variate shifts, which is a real problem if the estimates are being used for decision-making in a multi-step search graph that can lead to branches that are unlike what the model/estimator saw at training/calibration (if not altogether out-of-distribution). Here are a couple papers that describe how to approach those challenges:
[1] Similarity-Distance-Magnitude Activations. In Findings of the Association for Computational Linguistics: ACL 2026, pages 22037–22057, San Diego, California, United States. Association for Computational Linguistics.
[2] Introspectable, Updatable, and Uncertainty-aware Classification of Language Model Instruction-following. In Proceedings of the ACM Conference on AI and Agentic Systems (CAIS '26). Association for Computing Machinery, New York, NY, USA, 1259--1269.
am I missing something or is there an incredible amount of title editorialization here? on the page itself (and within the slug), the title is:
"Benchmarking AI decision models against traditional guardrails"
Agree. In fact Jev does reasonably well in both tests. In the article, the assertion is hedged, and followed up with a sentiment I liked:
> decision models like Jev do not reliably outperform [other] models in speed or accuracy. However, they rightly refocus industry attention on lightweight, task-specific inference paradigm that more closely resembles predictive machine learning
(I also noticed some issues with bolding in Table 4 that downplay Jev's performance a little.)
Jev (or equivalent) seem to me ideally placed for prototyping - prove the classification (say) works or is useful, then decide on how that should be made robust and economic with other options (eg local) brought in to the evaluation. Much quicker to get up and running, especially in greenfield scenarios.
Traditional classifier isn’t a direct equivalent though. Jev has some light abstraction/reasoning ability.
eg feed it a weather forecast and ask it whether I need an umbrella. It’s smart enough to make the connection between rain and umbrella.
So somewhere between classifier and fat LLM.
Ultimately boils down to right tool for the job
BART is quite an old model for this kind of test, and probably not a good very good choice for much these days. I'm working on replicating their benchmark on my own NLI model that targets zero-shot guardrail applications. I don't think it'll beat much larger models, but should give a better baseline for what a crossencoder can do.
https://huggingface.co/dleemiller/crossingguard-nli-l
Shouldn't LLMs intuitively be better with a high number of available options? This article only does simple prompts with only options to block or not block.
What is openai doing for their decisions API, a finetuned luna?
If you build a traditional classifier, and you have the data set for training curated or created hy an LLM, then you would essentially be building a classifier that judges the same way the LLM would?
There are plenty of benchmarks that show they do, too, though, so this is a single data point.
prompts that these evaluations were done are too trivial
The block/allow framing is the part I'd push on. I had a batch where automated checks passed all 99 outputs and reading each one by hand found 8 broken. The scorer only catches the failure modes its rubric already names, and a two-option guardrail bench inherits that ceiling whichever model sits behind it.