Science or Slop?: Benchmarking and Mitigating Scientific Slop in AI-Generated Papers
QUESTION — How can flawed reasoning patterns in AI-generated scientific papers be detected and mitigated?
This paper benchmarks scientific slop in AI-generated papers through six measures across Structure, Argument, and Artifacts, using a new dataset called SciSlopBench containing 390 AI-generated papers paired with human-written ones. The authors propose SciSlopHarness, a framework guiding LLMs to revise slop based strictly on experiment records, reducing the AI-human gap without requiring human reference targets.
Our measures identify the AI paper in each pair with 85.9% accuracy, compared with 68.7% for Binoculars.
SciSlopHarness reduces the remaining AI-human gap by 63% over the strongest revision baseline.
yerim0210 · 30 Sept 2026
read the original ↗