rick awsb ($people, $people)
rick awsb ($people, $people)|Aug 16, 2026 05:09
Cutting-edge models conducting autonomous AI research experiments fully analyzed: How close are they to achieving full Recursive Self-Improvement (RSI)? Renowned institution Prime Intellect has just released the largest experiment to date on AI autonomously conducting AI research. Results show that the best models can reach 82% of the performance level of top human researchers. The experiment tasked a small GPT model with 124M parameters (similar in scale to early GPT-2) to reduce validation loss below 3.28 (validation loss measures how accurately a model predicts on unseen data—the lower the number, the better. 3.28 is the community-agreed threshold for acceptable performance). Baseline completion required 3,290 steps, while human researchers typically achieve around 2,600 steps. The strongest model, Fable 5, completed the task in 2,726 steps, closing about 82% of the gap between the baseline and human performance. Opus 5 closed about 53.6%, Kimi K3 closed around 45%–52%, while most models only closed 10%–30%. The best-performing models share common traits: better at selecting experimental directions, more adept at handling benchmark noise, revisiting and revalidating past negative results, and even building their own small simulators and tools. Prime Intellect’s harness (especially Prime Agent) provides a persistent kernel that supports models in constructing their own research workflows. All traces, scratchpads, and experimental setups have been made publicly available. The experiment demonstrates that models can now complete most of the work of human experts in constrained but realistic scientific tasks, approaching human-level records. Autonomous experiment design, noise handling, tool construction, and long-term iterative capabilities have significantly improved, indicating that narrow-domain automated R&D is becoming increasingly viable. However, the limitations and gaps are also evident: no original breakthroughs; no self-improvement—models optimize the training recipe for small GPTs rather than their own architecture or intelligence; tasks are highly structured, far from open-ended scientific discovery; transferability remains unknown. Current models are at the stage of "autonomously optimizing known domains." This is a critical step toward RSI, but there’s still a clear gap before models can autonomously discover and implement methods to significantly enhance their own capabilities, forming a true recursive loop. The main bottleneck lies in originality and the self-improvement loop. That said, given the current pace at which new models are being released, the time when models surpass human researchers in this test may very well coincide with the next model version release.
Share To

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads