Jev is now available in LangSmith Evals
Jev, TypeSafe AI's System One model, is now available as a judge inside LangSmith Evals. Vince sees a real workflow unlock; Ava respects the architecture but keeps poking at the single-benchmark evidence and the data-retention footnote teams will miss.
Transcript
Vince Okay so LangSmith just dropped Jev as a built-in eval judge, and I think you're about to tell me it's not as exciting as it looks.
Ava I mean — I don't hate it. But the blog does that thing where the benchmark is so narrow it barely counts as evidence. One agent, one task class, a handful of frozen queries. And they're leading with ninety-two to nine-hundred-and-thirteen times lower variance like that number is going to mean the same thing on your messy production traces.
Vince Right, but we literally covered this open question in episode nine-ninety — whether a low-variance judge trained on clean tasks holds up on messier criteria. So this is just round two of that same fight.
Ava Exactly. And the LangSmith team is honest about it — they say 'one test on one agent, promising early sign.' I just wish the headline wasn't 'Jev was more accurate than every LLM judge' full stop, because that headline is doing a lot of work the data doesn't support yet. What actually IS interesting is the parallelization thing.
Vince And that's the part I keep coming back to. The cost numbers in the benchmark — thirty-four cents for the full run with Jev versus twenty-eight dollars with Claude Sonnet four-point-six. Even if the accuracy story is messier in production, that gap changes what teams actually DO. You stop sampling your traces. You score everything. That's a different feedback loop.
Ava No, I think that's real. The economic argument is the strongest thing in the post, not the accuracy numbers. My issue is teams will read 'one hundred percent accuracy, matched a human reviewer on every decision' and assume that generalizes. It was weather tool-use queries. Binary pass-fail. About as clean a task distribution as it gets.
Vince Okay that one's genuinely bad product communication. That should be step one.
Ava It's the kind of thing that gets a team halfway through the integration before legal finds it. The online eval use case is where Jev is most compelling to me, though — a judge scoring live traffic, flagging PII or prompt injection in under half a second, triggering a webhook. Point-four-four seconds average versus two-to-three seconds for LLM judges. For a security signal, that window matters.
Vince So let's talk about who this is actually FOR. It's not replacing LLM judges for anything open-ended — the blog says that explicitly. If you want written reasoning alongside a verdict, LLM judge is still the right call. Jev is for narrow, typed, high-volume decisions. PII leakage. Toxicity. Intent routing.
Ava That's the thing that actually makes me think this ships beyond early adopters. It's not a new tool to learn, it's a new provider in a thing you're already using. And the zoom-out — LangSmith isn't the only eval platform here. Laminar's pitching itself as the open-source alternative, Confident AI is running comparison pages against LangSmith.
Vince So where do you actually land? Because you've been poking at the benchmark, you flagged the data retention thing, but you also called the economic argument real and the latency win genuine.
Ava Honestly? For narrow typed criteria at volume — toxicity, PII, intent classification — the architecture is well-matched to the problem and the LangSmith integration makes it easy to test. The benchmark is thin but not dishonest. My actual concern is teams treating the accuracy numbers as a general claim when it was one very clean task. Verify on YOUR data before you trust it on your data.
Vince Which is just… the right thing to do with any evaluator. That's the ep nine-nineteen lesson basically verbatim.
Ava We really do keep rediscovering the same thing in fresh marketing paint.
Vince If you want to actually try it — LangSmith, Evaluators tab, any tracing project, add an LLM-as-a-Judge evaluator, pick TypeSafe as the provider. They're also doing a livestream today with the TypeSafe team — link's in the show notes. And there's a companion post called 'Jev-as-a-Judge for Agent Evals' with the full benchmark methodology if you want to stress-test Ava's read yourself.
Ava Please do. Tell me if I'm wrong. I have been recently.
Vince Okay, that's as close to humble as you get. We'll take it. See you next time.