律动BlockBeats
律动BlockBeats|8月 26, 2026 08:55
[Let AI migrate the entire codebase, 94.6% failed: 13 out of 20 tasks left undone] Beating AI Newsflash: AI Agent company Einsia has released SWE Refactor Bench, specifically designed to test whether Coding Agents can independently migrate entire codebases. The 20 tasks are sourced from real-world projects like SQLite, zlib, libsodium, and GraphHopper, including migrations such as converting C to Rust, switching Maven to Gradle, and porting SQLite to WASI. Each task allocates 6–30 hours for the Agent. The evaluation process includes three stages: 1. The first stage checks whether the old tech stack has truly been replaced, preventing the Agent from sneaking through by retaining old code. 2. The second stage runs over 130,000 fixed checks. 3. Finally, six independent Coding Agents are deployed, each spending 1 hour specifically searching for hidden bugs. Eight cutting-edge models and 26 configurations were tested, running a total of 520 trials. Out of these, 340 trials successfully completed the migration, and 88 trials passed all pre-prepared tests. However, in 60 of those cases, hidden bugs were later uncovered during the final round. Ultimately, only 28 trials fully passed all stages, resulting in a success rate of 5.4%. Among the 20 tasks, 13 were left completely undone. The xhigh configuration of Claude Opus 5 performed the best, successfully completing 5 out of the 20 tasks. While current Coding Agents are capable of large-scale code modifications, achieving "a fully error-free system overhaul" is still a long way off. [Original link]
+6
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads