律动BlockBeats
律动BlockBeats|Aug 05, 2026 10:43
[Qwen Matches Opus 5 in 3-Round Recall Rate, at Half the Cost] According to monitoring by Beating, the security company Aikido used its code auditing Agent to test seven models, including Qwen3.8-Max, Claude Opus 5, Kimi K3 Max, DeepSeek V4 Flash, as well as GPT-5.6 Sol, Luna, and Terra. The test involved 32 recently disclosed vulnerabilities, with each model running three times. Qwen3.8-Max identified a total of 26 vulnerabilities across three rounds, achieving a recall rate of 81.3%, tying for first place with Opus 5. Its F1 composite score (balancing false negatives and false positives) was 83.2%, slightly lower than Opus 5, Kimi K3 Max, and GPT-5.6 Sol. The issue with Qwen lies in its inconsistent performance across individual runs. Out of the 26 vulnerabilities, only 10 were consistently identified in all three rounds, whereas Opus 5 and Sol each identified 19. However, Qwen's total cost was approximately $821, only half that of Opus 5 and Sol, though still five times that of DeepSeek V4 Flash. [Original Link]
+2
Mentioned
Share To

Timeline

HotFlash

APP

X

Telegram

Facebook

Reddit

CopyLink

Hot Reads