律动BlockBeats
律动BlockBeats|Aug 01, 2026 03:11
**[Arena Launches AutoEval: AI Simulates Human Scoring, New Models Can Be Ranked in Just 1 Hour]** According to monitoring by Beating, the large model competition evaluation platform Arena has launched AutoEval, which uses reward models (scoring models that learn human preferences) to simulate user voting and quickly generate leaderboard scores for new models. After a new model is released, there is no need to wait several days to accumulate human votes—an estimated ranking can be obtained in less than 1 hour. These scores will be labeled as AutoEval and later corrected based on actual human voting results. This system is trained on millions of real user preference data points. During evaluations, AutoEval compares the response quality of different models and calculates scores according to Arena's existing leaderboard rules. Currently, it supports the evaluation of text, visual, image generation, and code models. Arena's backtesting shows that the correlation between AutoEval and subsequent human rankings exceeds 0.98. When the actual score difference between two models is greater than 10 points, its accuracy in determining the winner exceeds 90%; when the difference exceeds 15 points, accuracy reaches 100%. However, when the score difference between models is less than 5 points, it remains challenging to reliably distinguish between them. [Original Link]
+5
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads