Just Ask Jev: Reinforcement Learning for Calibrated Decisions as a Zero-Shot Detector of AI Alignment Failures
QUESTION — How can alignment failures in deployed language models be detected efficiently and cost-effectively without relying on expensive generative judges?
The paper introduces Jev, a model trained with reinforcement learning for calibrated decisions (RLCD) to serve as a zero-shot detector of AI alignment failures. Unlike generative judges that spend extensive decoding passes per criterion, Jev answers multiple typed questions about a single input with calibrated probabilities in one call. Evaluated on RLCDAlignBench spanning ten alignment failures and 44 benchmarks, Jev achieves a median AUROC of 0.886 zero-shot, outperforming supervised baselines while costing 63x less than LLM-judge scorers.
A single generic question reaches a median AUROC of 0.886 zero-shot and beats supervised baselines on most benchmarks.
Jev costs 63x less than LLM-judge scorers.
sumleo · 24 Sept 2026
read the original ↗