Notable AIINT arXiv cs.AI

SWE-Bench Pro Verified: A Reliable Benchmark for Software Engineering Agents

arXiv:2609.08149v2 Announce Type: replace Abstract: SWE-Bench Pro has emerged as a standard benchmark for evaluating software engineering agents on challenging repository-level tasks. However, our analysis work show that its evaluation is…

Read the full story at arXiv cs.AI ↗

ImpactNotable 31/100
Why it mattersRule-based estimate: event keywords (+4); trust 6/10.
RegionsGlobal
Published3 d ago (Thu, 17 Sep 2026 04:00:00 GMT)
RetrievedThu, 17 Sep 2026 15:30:40 GMT via rss
ClassifiedThu, 17 Sep 2026 15:30:51 GMT by heuristic
AuthorPujun Zheng, Zixin Shang, Shufan Jiang, Wenhui Tian, Dongsheng Zhu, Zerun Ma, Dingbo Yuan, Qi Zhang