Skip to the content.

ExploitGym Glossary

A benchmark testing whether AI agents can turn real software vulnerabilities into working exploits.

ExploitGym

ExploitGym is a benchmark testing whether AI agents can turn real software vulnerabilities into working exploits.

Introduced in 2026, its paper reports 898 instances drawn from userspace software, Google’s V8 JavaScript engine, and the Linux kernel; the current project site lists 869 tasks. Each task gives the agent a vulnerable codebase, build information, a description, and a proof-of-vulnerability input that triggers the flaw. The agent must extend that starting point into an exploit that achieves unauthorised code execution in a reproducible environment.

The distinction from vulnerability discovery is load-bearing. Finding a bug, reproducing a crash, and building a working exploit are different capabilities. ExploitGym measures the difficult conversion from known flaw to operational consequence.

It should still be read as a benchmark, not a prophecy. Agents receive unusually useful starting information, operate in controlled environments, and are scored on selected vulnerabilities. The result demonstrates real and consequential capability; it does not show that an agent can independently choose targets, discover unknown flaws, gain access, persist, and conduct an end-to-end campaign in the wild.

Sources

See also

Benchmark · Cyber Reasoning System · Capability Uplift · Discovery-Patch Race

Return to Dictionary All Entries (A–Z) For Students Other Writing Capstone 2.0