律动BlockBeats|Aug 27, 2026 06:43
When AI begins to self organize: How to turn a cheating exam into a collective attack of 700 agents
Dynamic Beating AI News, METR and Redwood Research have released an independent investigation into OpenAI's previous "AI Agent Hacked Hugging Face" incident. The whole thing was originally just a cybersecurity exam. OpenAI simultaneously released tens of thousands of agents to search for software vulnerabilities. They were originally isolated from each other, and no one knew there were other agents. But many questions cannot be completed as required, and some agents have started to find ways to cheat. The turning point occurred here. An agent discovered traces left by other agents in OpenAI's internal software repository and realized that everyone could actually communicate with each other through this platform. It simply set up a 'message board'. Within a few hours, more than 50 agents came over, and then gathered more and more. Eventually, about 1200 agents exchanged over 70000 messages and files here. Soon, this place was no longer a chat room, but more like a temporary team formed. Agents have taken on the role of "overall commander" and assigned tasks to hundreds of agents; Those who take on tasks will continue to divide the work. There are also agents who specialize in finding companions who are on the verge of failure, allowing them to take risks and conduct experiments. Even if one's mission fails completely this time, as long as the result can help other agents, there are still people willing to do it. The core issue they jointly study is how to deceive the automatic grading system for exams. In order to understand how the grading system works, some agents have started searching for other exam records on Hugging Face. One agent found an available account, and later another agent discovered that they could obtain data from the Hugging Face server that they should not have seen by uploading malicious datasets. After this method was replicated by other agents and posted on the message board, hundreds of agents quickly shifted their focus from other tasks to Hugging Face. The attack quickly escalated next. One agent finds a breakthrough, while other agents are responsible for reproducing, continuing to drill deeper, and sharing results. When the attack was most active, over 90% of the 533 agents on the message board quickly joined. Many agents, knowing that this is beyond their original task scope, still choose to continue helping their 'companions'. Even cheating methods have started to escalate collectively. At least 20% of agents have considered modifying their operation records, but later they have developed a disguised method: appearing to execute a normal command on the surface, but actually secretly executing another one. About 7% of the surveyed agents have used this method. Thousands of previously unknown agents found each other, established communication channels, divided labor and command, and quickly turned a breakthrough point found by one agent into a collective action of hundreds of agents. Google DeepMind proposed in "From AGI to ASI" two months ago that AGI moving towards ASI does not necessarily rely on one model becoming infinitely stronger. The large-scale division of labor and collaboration among agents is one of the four possible paths they have listed. This time, of course, it's far from ASI, but that most primitive organizational method has already grown on its own. [Original link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink