Imprint Reader: From Weight-Update Readout to Behavioral Intervention
The authors introduce the Imprint Reader, a model trained with Semantic Mount-and-Read Tuning (SMaRT) to decode frozen weight updates into natural-language descriptions. SMaRT mounts each update onto the Reader, eliciting descriptions via an anchor-free meta-query. Beyond free-form generation, the Reader offers a differentiable proxy between a target behavior and a candidate weight update, enabling direct behavioral intervention through MetaEdit without target-task training data.
On held-out updates, the joint Reader reaches judge-based Pass@100 of 2% for knowledge and 16% for behavior.
At a 0.5% pruning rate, Reader-guided selection raises measured harmful-prompt refusal from 57.9% to 64.1%.
Without target-task training data, MetaEdit increases BFCL Overall from 41.69% to 44.60%.