A Missing Piece for Trustworthy AI Reviewers: From Benchmarking Rhetorical Robustness to SciCore Review
This work defines Rhetorical Robustness as the stability of AI reviewers across content-preserving rewrites and their discrimination across papers. The authors introduce RobustReview, a benchmark comprising 1,260 manuscript versions evaluated across 30 reviewer configurations, revealing vulnerabilities to rhetorical variations and false robustness. To address this, they propose SciCore, a dual-branch reviewer that averages full-manuscript judgments with evaluations based on an extracted, structured science core. In comparisons using GPT-5.5, SciCore achieves a leading joint stability-discrimination profile while retaining competitive human alignment.
RobustReview contains 1,260 manuscript versions evaluating 30 distinct reviewer configurations.
SciCore achieves a leading joint stability-discrimination profile in primary GPT-5.5 comparisons.