CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 23 upvotes

BiasReducer: Adaptive Bias Mitigation for Reward Models

QUESTION — How can bias toward superficial attributes be mitigated in reward models without full retraining?

Reward models often favor superficial attributes like length or confidence over actual quality. Existing methods require costly retraining or fixed edits for known biases. This study proposes BiasReducer, a lightweight framework that edits only the linear reward head and selects relevant edits for each new dataset using a sparse autoencoder-style encoder. This approach improves reward-model robustness to superficial biases and transfers effectively to downstream tasks.

BiasReducer-M improves the three benchmarks by 8.3, 18.0, and 6.9 percentage points on average across five reward models.

Olivia121 · 26 Sept 2026 read the original ↗
↑