Apps & Tools
AI Bug-Bounty Agent Benchmark
A benchmark of nine AI agents against 100 real-world bug-bounty findings rebuilt as labs.
Playable
AI Bug-Bounty Agent Benchmark media is blocked
Allow external media to connect to the provider and play this content.
Description
A comprehensive security benchmark testing nine AI agents against 100 synthetic black-box web and API environments derived from real-world bug-bounty findings. The evaluation measures solve rates, verified proof rungs, and normalized API costs for models including Claude Opus 5 and Grok 4.6.
Descriptions, tags, and model credits may be AI-generated or inferred from public sources and can be incomplete or wrong. Learn more.