PowerStep: Memory-Efficient Adaptive Optimization via $\ell_p$-Norm Steepest Descent
arXiv:2605.10335v2 Announce Type: replace-cross Abstract: Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial memory overhead. We introduce PowerStep, a…
Read the full story at arXiv cs.AI ↗
ImpactNotable 31/100
Why it mattersRule-based estimate: event keywords (+4); trust 6/10.
RegionsGlobal
Published1 h ago (Wed, 30 Sep 2026 04:00:00 GMT)
RetrievedWed, 30 Sep 2026 04:00:56 GMT via rss
ClassifiedWed, 30 Sep 2026 04:01:19 GMT by heuristic
AuthorYao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian