CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 13 upvotes

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation

QUESTION — How can backdoors be effectively removed from Large Language Models without degrading benign performance or safety?

The authors propose NEEDLE, a training-free method for targeted backdoor removal in LLMs. By estimating a backdoor direction and a refusal subspace through activation vectors, the method applies sequential weight orthogonalisation to suppress the backdoor while preserving refusal-related representations. Evaluations across multiple model families and attack types demonstrate that NEEDLE achieves the lowest mean Attack Success Rate and minimal changes in capability and safety.

NEEDLE is a training-free method for targeted backdoor removal that requires neither a clean reference model nor original poisoned data.

NEEDLE achieves the lowest mean Attack Success Rate among evaluated defences across multiple model families.

NEEDLE achieves 0% Attack Success Rate on challenging code injection attacks.

GeorgeDrayson · 29 Sept 2026 read the original ↗
↑