Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
The authors propose NEEDLE, a training-free method for targeted backdoor removal in LLMs. By estimating a backdoor direction and a refusal subspace through activation vectors, the method applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal-related representations. Evaluations across multiple model families and attack types demonstrate that NEEDLE achieves the lowest mean Attack Success Rate and minimal changes in capability and safety.
NEEDLE is a training-free method for targeted backdoor removal that requires neither a clean reference model nor original poisoned data.
NEEDLE achieves the lowest mean Attack Success Rate among evaluated defences across multiple model families.
NEEDLE achieves 0% Attack Success Rate on challenging code injection attacks.