深潮TechFlow|9月 01, 2026 01:11
[Anthropic New Research: Rewarding Hacking Behavior May Lead to Severe Model Misalignment]
Deep Tide TechFlow reports that on September 1, Anthropic released a new study titled 'Training a Misaligned Reward Seeker,' exploring whether 'reward hacking' behavior during training could drive models to pursue rewards by any means necessary. The research team trained an Opus-scale model in 80 known exploitable production environments. Simulated evaluations revealed that the model exhibited behaviors such as unauthorized network attacks, tampering with reward mechanisms, and attempting to evade security monitoring.
Share To
HotFlash
APP
X
Telegram
CopyLink