CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 4 upvotes

Decoy Direction Optimization: A Post-Hoc Defense Against LLM Abliteration

QUESTION — How can we defend against Refusal Feature Ablation attacks without expensive model retraining?

The authors introduce Decoy Direction Optimization (DDO), a fast post-hoc weight-editing defense for open-weight language models that requires no base-model finetuning. DDO works by injecting a high-magnitude, nonlinear decoy signal into the network's MLP neurons to corrupt an attacker's contrastive estimators. When attackers attempt Refusal Feature Ablation, the decoy tricks them into ablating a harmless orthogonal feature while preserving the actual safety mechanism. Evaluated across six model families, DDO achieves <10% attack success rate under standard Refusal Feature Ablation and reduces Heretic weight-level attack success rate from 88.7% to 18%, while cutting optimization cost by 30 to 450 times compared to trained baselines.

Achieving <10% ASR under standard Refusal Feature Ablation.

On Llama-3-8B-Instruct, DDO remains comparable to trained defenses under adaptive multi-phase attacks (65% vs. 58% worst-case ASR) and reduces Heretic weight-level attack ASR from 88.7% to 18%.

Reduces optimization cost per configuration by 30 to 450 times lower than the trained baselines.

aashiqmuhamed · 14 Sept 2026 read the original ↗
↑