When AI Reviews Train AI Reviewers: Scientific-Judgment Collapse and Mitigation
QUESTION — How does scientific-judgment collapse manifest when AI peer reviewers are recursively trained on synthetic reviews generated by earlier models?
This paper investigates the recursive feedback loop in AI scientific peer review, where successor reviewer models are trained on reviews generated by predecessor models. Starting from Llama 3.1 8B, the authors fine-tune a reviewer on official ICLR reviews and train successor models on systematically varied mixtures of official and model-generated reviews. The study shows that introducing synthetic reviews compresses rating distributions and reduces semantic diversity, a pattern termed scientific-judgment collapse. To mitigate this failure mode, the authors introduce TrustReviewer, an open-source LLM-based system that intervenes at training time via a curated training corpus and at test time via paired activation steering.
taesiri · 17 Sept 2026
read the original ↗