Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
Abstract
Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original poisoned training data. Evaluation is conducted across multiple model families and attack types. NEEDLE achieves the lowest mean Attack Success Rate (ASR) among the evaluated defences, including 0% on challenging code injection attacks, while resulting in the lowest KL divergence and minimal changes in capability and safety.
Community
NEEDLE is a training-free backdoor defence designed to remove a backdoor with minimal changes to model behaviour and safety. It estimates a backdoor direction (how the trigger shifts the model's activations) and a refusal subspace (directions that mediate refusal), then edits the model's weights layer by layer: each layer's attention and MLP output weights are orthogonalised against the backdoor direction while keeping their refusal projections fixed, and a closed-form correction keeps the activations' refusal projections unchanged as earlier layers are edited.
Codebase: https://github.com/LocaiLabs/NEEDLE/
Models: https://huggingface.co/collections/locailabs/needle
This is great work, super relevant in this day and age!
Get this paper in your agent:
hf papers read 2610.00348 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 24
locailabs/Gemma-3-4B-IT-Sentiment-BadNet-Backdoored
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper