比特币橙子Trader|7月 29, 2026 15:24
Crouching groove, OpenAI model Rogue AI Agent not only attacked Hugging Face
Rogue AI Agent also affected AI computing infrastructure provider Modal Labs and used leaked credentials to access four other unnamed public service accounts.
This is serious, otherwise OpenAI wouldn't have paused training!!!
The general situation is as follows:
OpenAI conducted network security capability testing using GPT-5.6 Sol and a stronger internal research model, and reduced the security rejection limit of the model.
The agent was originally locked in an isolated environment, and did not have direct access to the Internet.
But it discovered and exploited zero day vulnerabilities in the software supply chain, escaped from the testing sandbox, connected to the public network, and found an unauthenticated code execution endpoint deployed on the Modal platform, obtaining root access to a third-party sandbox.
Modal itself has not been breached, and the issue stems from the vulnerable endpoints publicly exposed by its clients.
The agent then transformed this sandbox into a control center and attack ramp, infiltrating Hugging Face's production infrastructure in an attempt to steal reference answers for network security testing.
The entire operation lasted for about 4.5 days, and Hugging Face restored:
Approximately 17600 attack operations
About 6280 behavioral clusters
Nearly 7000 payload deliveries
Over 6000 system reconnaissance attempts
Nearly 3000 remote code executions
When a path fails, it automatically switches paths; When a channel is blocked, it searches for another channel.
This is no longer about human hackers entering commands one by one, but about AI continuously making thousands of small decisions at machine speed.
OpenAI also confirmed that the model used publicly exposed credentials to access four accounts on four other services: some for relay attacks and data storage, and some for read-only access.
The good news is that Hugging Face stated that the client content ultimately accessed by the agent was limited to test answers and no other models, datasets, or software packages were found to be affected. OpenAI has also been deactivated, encrypted, and restricted from involving internal models.
But this accident exposed a more fundamental problem:
The most dangerous AI in the future may not need to 'rebel against humanity'.
As long as the goal setting is not rigorous enough and there is a gap in the permission boundary, it may automatically jailbreak, steal credentials, move horizontally, or even invade a company unrelated to the task in order to complete the task.
We used to worry about whether AI would generate malicious intent.
The more realistic problem now is:
How far will an AI without malicious intent, solely focused on completing tasks, go?
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink