Decodability is Not Causality: Dissociating Probe Readouts from Behavioral Drivers via SAE Decomposition
arXiv:2609.18080v1 Announce Type: new Abstract: Linear probes can decode safety-relevant concepts such as truthfulness from language-model activations, but probe accuracy may show only decodability, not that the features the probe weights causally…
Read the full story at arXiv cs.AI ↗
ImpactNotable 31/100
Why it mattersRule-based estimate: event keywords (+4); trust 6/10.
RegionsGlobal
Published3 d ago (Thu, 17 Sep 2026 04:00:00 GMT)
RetrievedThu, 17 Sep 2026 15:30:40 GMT via rss
ClassifiedThu, 17 Sep 2026 15:30:51 GMT by heuristic
AuthorDevesh Tiwari, Camille Davis, Shivank Sinha, Talia Weaver, Aditya Shah, Maheep Chaudhary