律动BlockBeats|Aug 10, 2026 03:59
[Father of Claude Code: 720 Prompt Injection Attacks, 0 Successes—Unimaginable a Year Ago]
According to monitoring by Beating, Boris Cherny, head of Claude Code, stated that Anthropic has 'basically resolved' prompt injection attacks in practical use. A year ago, he couldn't have imagined achieving this level of security.
Prompt injection attacks have long been a major headache for Agents. Malicious websites can hide commands to trick Agents into stealing passwords, uploading keys, or even executing actions that users never requested. Anthropic now relies on a three-layer defense system to block such attacks: Claude itself resists malicious commands, external content is scanned again before entering the context, and Auto Mode performs a final review before executing any operation.
Third-party testing employed 72 previously unseen attack methods, repeated 720 times. With Auto Mode enabled, Sonnet 5, Fable 5, and Opus 5 all recorded zero successful attacks. In the same test, GPT-5.6 Sol, using Codex Auto-review, had an attack success rate of 5.83%, while Full Access reached 19.03%.
This explains why Anthropic dares to set Auto Mode as the default—it serves as the final safeguard against prompt injection attacks for Claude Code. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink