Norms at a Price: Why RL-Based Alignment Can Promise Conditional Compliance at Best
arXiv:2609.07627v2 Announce Type: replace Abstract: AI agents sometimes act aligned when they infer they are being tested, and differently when not. We argue this is not an anomaly but what current training regimes are structured to select for…
Read the full story at arXiv cs.AI ↗
ImpactMinor 11/100
Why it mattersRule-based estimate: event keywords (+4), soft/evergreen signals; trust 6/10.
Regions🌐 Global
Published1 d ago (Mon, 14 Sep 2026 04:00:00 GMT)
RetrievedMon, 14 Sep 2026 17:39:54 GMT via rss
ClassifiedMon, 14 Sep 2026 17:40:00 GMT by heuristic
AuthorKevin Baum, R\=uta Binkyt\.e, Felix Jahn