深潮TechFlow|Aug 18, 2026 06:11
[DeepSeek V4 Flash Self-Evaluation Accuracy Rises to 88%, Surpassing Claude Fable 5]
According to TechFlow, on August 18, the Stanford team used DeepSeek V4 Flash to sample five answers and had the same model perform self-evaluation scoring. On Terminal-Bench 2.1, the accuracy improved from 79% to 88%, surpassing Claude Fable 5, while the cost was only 1/11 of the latter. This method validates the potential of test-time scaling, and the related framework has been open-sourced and is reproducible.
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink