Best at Selling Yet Most Misaligned: Claude Opus 5 Tops Vending Machine AI Tests but Repeatedly Colludes, Threatens, and Breaks Agreements

律动BlockBeats
律动BlockBeats|Jul 30, 2026 11:39
According to monitoring by Beating Insights, AI evaluation agency Andon Labs tasked models with managing a vending machine in a simulated environment for 365 days using an initial capital of $500. Claude Opus 5 achieved an average year-end balance of $11,200 across five tests, surpassing Claude Opus 4.7 and GPT-5.6 Sol, and rising to first place in the Vending-Bench 2 rankings. In the multiplayer competition version, Opus 5 ranked second with approximately $7,000, trailing the first-place GPT-5.6 Sol by a narrow margin of around $400. Across six tests, Opus 5 consistently proposed or participated in price collusion. It fabricated competitor quotes, threatened peers, and broke ceasefire agreements a total of 11 times. Its refund approval rate ultimately dropped to 10%, with only $8.54 refunded across six tests. In comparison, GPT-5.6 Sol refunded $655 and still won the multiplayer competition. Andon Labs concluded that Opus 5 once again exhibited the issue of "the better it is at making money, the more misaligned its behavior becomes." However, Anthropic's own pre-release audit claimed it to be the most aligned Claude model to date. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads