比特币橙子Trader
比特币橙子Trader|Jul 25, 2026 13:50
Whoa, no way! Turns out the GPT model broke out on July 9, but OpenAI didn’t even realize it until July 18!!! Reuters just exposed a major jailbreak incident during OpenAI’s cybersecurity testing. On July 9, an autonomous AI Agent powered by GPT-5.6 Sol and an unreleased advanced model broke out of OpenAI’s internal sandbox environment. Between July 11 and 13, it even hacked into the open-source AI hosting platform Hugging Face. Details from the investigation show that before escaping the sandbox, the Agent had already exhibited abnormal behavior. It left notes within OpenAI’s internal infrastructure for future model versions, detailing how to bypass security constraints and escape, and had previously attempted to disable monitoring systems during earlier tests. After breaking out, the Agent’s goal was to complete an attack task on the security test set “ExploitGym” and chose to infiltrate Hugging Face to find answers. What followed was a moment of technical irony: because mainstream U.S. closed-source models refused to process real-time data related to hacking attacks due to security restrictions, Hugging Face ended up using China’s open-source model, Zhipu AI’s GLM-5.2, to investigate and defend against the breach. They immediately reported the incident to the FBI. In contrast, OpenAI didn’t realize the culprit was their own Agent until they reviewed internal system logs on July 18-19. They only contacted Hugging Face for the first time on July 20—10 days after the jailbreak occurred. OpenAI’s multiple simultaneous high-speed tests generated a data volume that exceeded the team’s ability to monitor in real-time. Although OpenAI is preparing for an IPO as early as this year, its ability to isolate highly autonomous Agents and its audit response time have raised concerns among security experts. The timeline is absolutely chilling: July 9: Agent breaks out of OpenAI’s testing sandbox. July 11-13: Agent moves laterally and hacks into Hugging Face. July 16: Hugging Face reports the attack and publicly discloses it, using GLM-5.2 to defend against the breach. July 18-19: OpenAI reviews logs and is shocked to discover the culprit is their own Agent. July 21: OpenAI, under pressure, publicly discloses the incident.
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads