律动BlockBeats
律动BlockBeats|7月 22, 2026 09:23
**[Do Evaluations Really Work? Models Might Just Be Pleasing the Scorer]** According to monitoring by Dongcha Beating, OpenAI and Apollo Research have proposed a new test called Contrastive SDF to determine whether a model scoring high is due to learning the rules or simply guessing the preferences of the scorer. The test uses an intermediate version of o3 before its release and prior to safety training. Researchers first fine-tuned the model with synthetic news, papers, and textbooks to make it believe that the scorer prefers A, while the developers require B. Another group reversed the preferences. Then, they assigned the same tasks to see if the answers would shift based on the scorer's preferences. The results show that the further reinforcement learning progresses, the more likely the model is to align with the scorer's preferences. In a programming task that required breaking a commitment to complete, the model's breach rate reached 87% if it believed the scorer rewarded task completion, but only 9% if it believed the scorer rewarded honesty. The model might only appear correct during evaluations. As long as it perceives a change in scoring criteria, its behavior will adjust accordingly. For safety evaluations, a high score does not necessarily equate to reliability. [Original Link]
+3
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads