The Unequal Influence of Bad Advice: Using Training Data Attribution to Modulate Emergent Misalignment
arXiv:2609.37914v1 Announce Type: cross Abstract: Fine-tuning large language models on narrow, misaligned tasks can undo their post-training alignment and induce novel misaligned behaviors -- a phenomenon known as \emph{emergent misalignment} (EM)…
Read the full story at arXiv cs.AI ↗
ImpactNotable 31/100
Why it mattersRule-based estimate: event keywords (+4); trust 6/10.
RegionsGlobal
Published1 h ago (Wed, 30 Sep 2026 04:00:00 GMT)
RetrievedWed, 30 Sep 2026 04:00:56 GMT via rss
ClassifiedWed, 30 Sep 2026 04:01:19 GMT by heuristic
AuthorGon\c{c}alo Paulo, Louis Jaburi, Nora Belrose, Lucia Quirke, Stella Biderman