律动BlockBeats|8月 06, 2026 10:31
[Qwen3.8 Max Matches Opus 4.8 After Retesting, Still Trails Kimi K3]
According to monitoring by Beating and evaluation agency Artificial Analysis, after retesting Qwen3.8 Max, its Intelligence Index was raised from 53 to 56. Previously, intermittent anomalies occurred in the testing interface, but this time the official Alibaba Cloud API was used for reruns. The new score matches Claude Opus 4.8 but remains 1 point lower than Kimi K3. The most significant improvement for Qwen is in Agent tasks. Its GDPval-AA score reached 1739, surpassing Kimi K3's 1685 and GPT-5.6 Sol's 1730, trailing only Claude Opus 5. Terminal operations, scientific reasoning, and programming evaluations also showed general increases. The trade-off is higher Token consumption. While Qwen3.8 Max's Token unit price is lower than the previous generation, the average cost of completing a comprehensive evaluation task still rose from $0.53 to $1.14, higher than Kimi K3's $0.86. The main reason is its increased rounds of calls in Agent tasks, generating more content. Its knowledge reliability has also regressed. AA-Omniscience accuracy remained stable, but hallucination rates increased from 23% in the previous generation to 40%. [Original Link]
Share To
Timeline
HotFlash
APP
X
Telegram
CopyLink