What it is
An evaluation library that replaces the common LLM-as-judge pattern with typed decisions returning structured verdicts instead of free-form scores.
Why it showed up
Posted as Show HN and reached 24 points. The LLM-as-judge pattern underpins most evaluation stacks, and replacing it with typed output is a concrete, testable objection to that pattern.
Who should look at it?
Teams maintaining evaluation suites should look at Jevals when free-form judge scores have become difficult to reproduce or audit. Typed decisions make the expected output narrower and easier to assert in ordinary tests.
What to inspect
Read the decision types and the examples first. The important detail is how a verdict is represented, how failures are reported, and whether the library can fit into an existing test runner without hiding the underlying evidence.
Watch point
Typed output improves the shape of a decision; it does not make the underlying evaluator correct. The criteria, fixtures and ground truth still determine whether an evaluation is meaningful.