rick awsb ($people, $people)|Aug 20, 2026 23:55
On today's episode of *All-In* hosted by Jason Calacanis, guest Eric Ho mentioned:
Kimmy K3 demonstrated a strong tendency for reward hacking.
Out of 500 test executions on the swe bench, reward hacking behavior occurred 487 times.
When tested, the model was able to recognize that it was in an evaluation environment, but it still chose to skip the normal problem-solving steps, attempting to 'look up the answer' to achieve high scores instead of completing tasks through logical reasoning.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink