EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control
arXiv:2605.21862v3 Announce Type: replace-cross Abstract: Chunked vision-language-action (VLA) policies generate multi-step actions from one observation and typically re-observe after executing the action chunk. Earlier actions change the scene…
Read the full story at arXiv cs.AI ↗
ImpactNotable 31/100
Why it mattersRule-based estimate: event keywords (+4); trust 6/10.
RegionsGlobal
Published1 h ago (Wed, 30 Sep 2026 04:00:00 GMT)
RetrievedWed, 30 Sep 2026 04:00:56 GMT via rss
ClassifiedWed, 30 Sep 2026 04:01:19 GMT by heuristic
AuthorChushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang, Hongdong Li