JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
The study investigates whether a decision-only judge (JEV) can provide an economical first-pass evaluation and determine when stronger evaluation is required. Comparing JEV against sixteen generative and reward-model judges with blinded human adjudication, JEV is within three percentage points of a state-of-the-art LLM judge on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee. Larger performance gaps occur when checking derivations or resisting elaborately written wrong answers. A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at a reduced cost.
JEV is within three percentage points of a state-of-the-art LLM judge on ordinary preference and evidence-grounded factuality at 0.36% of the comparator's fee.
A frozen cascade that accepts confident verdicts and escalates uncertain ones retains 99% of the comparator's accuracy at lower cost.