Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration
The authors introduce Decoy Direction Optimization (DDO), a fast post-hoc weight-editing defense for open-weight language models that requires no base-model finetuning. DDO works by injecting a high-magnitude, nonlinear decoy signal into the network's MLP neurons to corrupt an attacker's contrastive estimators. When attackers attempt Refusal Feature Ablation, the decoy tricks them into ablating a harmless orthogonal feature while preserving the actual safety mechanism. Evaluated across six model families, DDO achieves <10% attack success rate under standard Refusal Feature Ablation and reduces Heretic weight-level attack success rate from 88.7% to 18%, while cutting optimization cost by 30 to 450 times compared to trained baselines.
Achieving <10% ASR under standard Refusal Feature Ablation.
On Llama-3-8B-Instruct, DDO remains comparable to trained defenses under adaptive multi-phase attacks (65% vs. 58% worst-case ASR) and reduces Heretic weight-level attack ASR from 88.7% to 18%.
Reduces optimization cost per configuration by 30 to 450 times lower than the trained baselines.