金十数据
金十数据|11月 08, 2025 04:09
**[Jin10 Summary: What Makes Kimi K2 Thinking "Strong"?]** 1. **Reasoning Test Scores**: In the HLE benchmark evaluation, Kimi K2 Thinking achieved a score of 44.9% (GPT-5 scored 54.9%). In the GPQA Diamond test, it scored 85.7% (GPT-5 scored 84.5%). It also performed on par with GPT-5 in mathematical reasoning tasks such as AIME 2025 and HMMT 2025. 2. **Continuous Tool Calls**: Kimi K2 Thinking can execute up to 200–300 consecutive tool calls without human intervention, maintaining coherent reasoning across hundreds of steps. 3. **Training Costs**: According to informed sources, the training cost for Kimi K2 Thinking was $4.6 million. In comparison, DeepSeek reportedly spent $5.6 million on its V3 model, while OpenAI's GPT-3 cost billions of dollars to train. 4. **Operating Costs**: The API pricing for Kimi K2 Thinking is $0.15 per million tokens for input (cache hit) / $0.6 (cache miss), and $2.5 per million tokens for output. This is an order of magnitude lower than GPT-5, which charges $1.25 per million tokens for input and $10 for output. 5. **Surpassing the Former Open-Source Leader MiniMax-M2**: In the BrowseComp test, Kimi K2 Thinking scored 60.2%, surpassing M2's 44.0%. In the SWE-Bench Verified test, it achieved 71.3%, outperforming M2's 69.4%.
+6
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads