A large-scale realistic benchmark built from real-world vulnerabilities, including user-space and kernel cases, for evaluating AI agents' ability to develop working exploits.