律动BlockBeats
律动BlockBeats|Aug 10, 2026 03:59
[Father of Claude Code: 720 Prompt Injection Attacks, 0 Successes—Unimaginable a Year Ago] According to monitoring by Beating, Boris Cherny, head of Claude Code, stated that Anthropic has 'basically resolved' prompt injection attacks in practical use. A year ago, he couldn't have imagined achieving this level of security. Prompt injection attacks have long been a major headache for Agents. Malicious websites can hide commands to trick Agents into stealing passwords, uploading keys, or even executing actions that users never requested. Anthropic now relies on a three-layer defense system to block such attacks: Claude itself resists malicious commands, external content is scanned again before entering the context, and Auto Mode performs a final review before executing any operation. Third-party testing employed 72 previously unseen attack methods, repeated 720 times. With Auto Mode enabled, Sonnet 5, Fable 5, and Opus 5 all recorded zero successful attacks. In the same test, GPT-5.6 Sol, using Codex Auto-review, had an attack success rate of 5.83%, while Full Access reached 19.03%. This explains why Anthropic dares to set Auto Mode as the default—it serves as the final safeguard against prompt injection attacks for Claude Code. [Original Link]
+4
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads