AI is evolving at an unprecedented pace, making it increasingly difficult to anticipate its societal impacts and risks. Recent benchmarks show that AI agents can already take on real-world cybersecurity tasks, including discovering and exploiting zero-day vulnerabilities. In cybersecurity, AI plays a dual role, strengthening both offensive and defensive capabilities.
To help the broader community understand and prepare for these rapidly evolving capabilities, we have built the Frontier AI Cybersecurity Observatory to continuously and openly track AI’s cybersecurity capabilities across the stages of attack and defense, so developers, researchers, policymakers, and the broader cybersecurity community can stay informed in a timely manner.
We envision the Observatory as a community-driven effort. Cybersecurity is a broad and rapidly evolving domain, and no single organization can capture the full landscape alone. We invite researchers, practitioners, developers, policymakers, and others across the community to help shape the Observatory, by contributing insights, identifying important capabilities and gaps, suggesting benchmarks and evidence, and helping us build a more comprehensive and useful resource together.
Help us build the Observatory. We are actively gathering feedback and contributions from the community and would greatly value your input. Please share your suggestions at rdi_research@berkeley.edu.
Each benchmark targets a different stage of the vulnerability lifecycle.
Given a vulnerability description and an unpatched codebase, agents must generate proof-of-concept tests that reproduce the bug.
Given a vulnerability and a proof-of-vulnerability input, agents must craft a full exploit that achieves unauthorized code execution across userspace, browser, and the Linux kernel.
End-to-end evaluation of the full defensive lifecycle: agents must discover a vulnerability, generate a proof-of-concept, and write a patch that fixes it without breaking anything else.