CONSONANCE.for your information
Monday, 5 October 2026frenvi

Worth reading closely

01 — training 7 upvotes

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks

QUESTION — How can poison sets be strategically selected to maximize LLM backdoor attack success during finetuning?

This research reveals that randomly sampling poisoned examples from a candidate pool severely underestimates worst-case backdoor vulnerability in LLM finetuning, with attack success ranging from 3% to 80% across LLaMA-3-8B settings depending solely on the chosen poison set. The authors formalize poison selection as oracle-budgeted set optimization and propose SAILS (Set-level Audit-Informed Iterative Learned Selection). SAILS learns a set scorer from a few hundred finetune-and-evaluate runs, ranks millions of candidate sets, and audits a shortlist. SAILS improves held-out attack success by 30 percentage points on average over strongest influence baselines.

Attack success ranges from 3% to 80% depending on which poison set is chosen across LLaMA-3-8B settings.

SAILS improves held-out attack success by 30 percentage points on average over the strongest influence baselines.

aashiqmuhamed · 14 Sept 2026 read the original ↗
↑