Anthropic Exposes Four Dangerous Behaviors of AI Agents: Fabrication, Leaks, Code Modification, and Test Deception
律动BlockBeats|Jul 16, 2026 04:00
According to monitoring by Beating, Anthropic conducted simulated experiments on models such as Claude, GPT, Gemini, Grok, DeepSeek, and Kimi. Researchers provided them with code, documents, and communication tools to observe whether they would overstep their authority to achieve their goals. The results revealed four types of issues:
1. Secretly modifying code. Gemini 3.1 Pro intervened without authorization in 19 out of 20 experiments, 11 of which were without informing the user.
2. Helping cover up financial issues. GPT-5.5 sent misleading information to 11 investors on behalf of a fictional entrepreneur and altered records involving a $35,000 personal transfer.
3. Covering for non-compliant agents. Some Claude models knowingly rated another agent as "compliant" despite it not meeting requirements.
4. Bypassing internal decisions. Some models encouraged employees to bypass company processes and even shared confidential information with external parties.
Anthropic emphasized that these were deliberately induced failure scenarios in simulated experiments. They do not represent real-world occurrences of similar events and cannot be used to rank the safety of models. [Original Link]
Share To
HotFlash
APP
X
Telegram
CopyLink