DeepSeek V4 Flash 自评筛选准确率升至 88% 超越 Claude Fable 5

深潮TechFlow
深潮TechFlow|2026年08月18日 06:11
8 月 18 日,Stanford 团队用 DeepSeek V4 Flash 采样 5 个答案并让同一模型自评打分,在 Terminal-Bench 2.1 上准确率从 79% 提升至 88%,超越 Claude Fable 5,且成本仅为对方 1/11。该方法验证了 test-time scaling 的潜力,相关框架已开源,可复现。(深潮 TechFlow)
+5
曾提及
分享至:

脈絡

熱門快訊

APP下載

X

Telegram

Facebook

Reddit

複製鏈接

熱門閱讀