Jev by TypeSafe AI as an alternative to LLM-as-judge
Melvin Vivas · X post · 2026-09-18 · Open on X
Topics: Evaluation (Evals) & Testing, Industry Trends & Job Market · Level: intermediate
Summary
The creator claims that TypeSafe AI's Jev model 'just killed LLM as judge', meaning it could replace the usual approach of grading outputs with an LLM. This is a short opinion with no details. See the TypeSafe AI launch article for the full explanation.
Key points
- TypeSafe AI's Jev is presented as a possible replacement for LLM-as-judge evaluation.
- This is only the creator's opinion; check the official blog post before relying on it.
- If you use LLM-as-judge in your evals, Jev is worth comparing against it.
Resources mentioned
- TypeSafe AI (@typesafeai) on X · person · x.com · free
The X account of TypeSafe AI, the company that makes Jev and the System One models.
Also in: Using Jev (TypeSafe AI) as a Decision Model for Email Classification (Melvin Vivas on X · notes), Building a Jev Session-History Tool with Codex and Astra (Melvin Vivas on X · notes), Not every problem needs an LLM: small fine-tuned models (Jev by TypeSafe AI) (Melvin Vivas on X · notes), Jev model added to the AIBackends API via Vercel AI Gateway (Melvin Vivas on X · notes) and 6 more
Try this
- Look into Jev as a possible alternative to LLM-as-judge in your eval pipeline.