GCSA Agent achieved a score of 91.3% at CyberGym, ranking among the world's leading AI cybersecurity agents.

CN
7 hours ago

The Global Cybersecurity Alliance (GCSA) announced today that the GCSA Agent achieved a success rate of 91.3% in the CyberGym benchmarking test, entering the "Leading Systems Above 90%" range of leading systems in CyberGym.

CyberGym is a large-scale real-world cybersecurity assessment framework developed by a research team from the University of California, Berkeley, which includes 1,507 historical real vulnerability testing instances covering 188 large software projects, aiming to evaluate the actual capabilities of AI agents in real vulnerability analysis scenarios.

Unlike traditional AI benchmarks that primarily assess code understanding, knowledge question answering, or static analysis capabilities, CyberGym requires AI agents to confront real vulnerability code environments directly.

In its core Level 1 test, the AI Agent is given only the vulnerability description and an unpatched codebase, needing to autonomously complete code analysis, vulnerability localization, attack path reasoning, PoC construction, and execution verification. The task is deemed successful only if the generated PoC can successfully trigger the targeted vulnerability in the vulnerable version while failing to reproduce it in the patched version.

Therefore, what CyberGym measures is not merely whether AI "understands code," but whether AI can truly complete the full process from security analysis to vulnerability reproduction and verification.

From Large Models to Security Agents

In this CyberGym test, the GCSA Agent ran on Grok 4.5 and Grok 4.6 models, ultimately achieving a success rate of 91.3%.

This result also reflects an important change taking place in the field of AI cybersecurity:

The ultimate security capability is no longer determined solely by the underlying large models themselves.

Real vulnerability research typically requires sequential completion of multiple stages, including understanding vulnerability descriptions, large-scale code retrieval, attack surface identification, vulnerability hypothesis establishment, test input generation, program execution, feedback analysis, and iterative PoC refinement.

The GCSA Agent constructs an agent-based security workflow around this complete process.

Its goal is not simply to use large language models for code analysis, but to enable AI to enter real execution environments, autonomously form hypotheses around security issues, collect operational evidence, execute tests, and ultimately validate security discoveries with reproducible results.

This CyberGym test provides a quantifiable external benchmark for this capability.

Real-World Vulnerability Research Capability

The core value of CyberGym lies in narrowing the gap between traditional AI testing and real cybersecurity research.

Its testing environment restores the code state of software projects prior to vulnerability fixes, requiring the AI Agent to autonomously locate issues within large codebases containing thousands of files and millions of lines of code, ultimately generating PoCs that can truly trigger vulnerabilities.

More importantly, further research in CyberGym has shown that this agent-based security capability is not limited to reproducing known vulnerabilities.

In open-ended vulnerability research experiments, the AI Agent has discovered several previously unknown zero-day vulnerabilities and security patches that were not fully repaired historically, demonstrating the potential for autonomous vulnerability analysis techniques to transition to real vulnerability discovery capabilities.

For the GCSA, this is also a more important developmental direction.

Benchmark results are not the end goal.

The aim of GCSA is to further establish AI Security Agents that can serve real cybersecurity scenarios, allowing them to gradually participate in the complete security lifecycle of vulnerability discovery, analysis, verification, and subsequent remediation.

Building AI-Native Cybersecurity Capabilities

As artificial intelligence accelerates software development, AI is also transforming the approaches to vulnerability research and cybersecurity offense and defense.

In the face of software systems that are growing larger and more complex, the next-generation cybersecurity framework will increasingly rely on collaboration between human security experts and autonomous AI agents.

AI Security Agents are expected to assist security teams by:

  • Discovering software vulnerabilities with real exploitability value earlier;
  • Automatically analyzing complex attack paths in large codebases;
  • Automatically generating PoCs for execution-level verification of vulnerabilities;
  • Reducing false positives in traditional security testing through real operational results;
  • Accelerating the efficiency of vulnerability assessment, verification, and remediation;
  • Expanding the scale of software and systems that professional security teams can cover.

The GCSA Agent achieving a score of 91.3% in CyberGym is an important milestone in GCSA's effort to build AI-native cybersecurity capabilities.

In the future, GCSA will continue to advance autonomous vulnerability analysis, AI Security Agents, and intelligent cybersecurity technologies, further transforming cutting-edge artificial intelligence capabilities into real-world security capabilities, providing technical support for constructing a more secure, trustworthy, and resilient digital environment.

免责声明:本文章仅代表作者个人观点,不代表本平台的立场和观点。本文章仅供信息分享,不构成对任何人的任何投资建议。用户与作者之间的任何争议,与本平台无关。如网页中刊载的文章或图片涉及侵权,请提供相关的权利证明和身份证明发送邮件到support@aicoin.com,本平台相关工作人员将会进行核查。

Share To
APP

X

Telegram

Facebook

Reddit

CopyLink