discovered 03 Aug 2026
ethibench
→ View on GitHubEthiBench is an adaptable evaluation framework designed for assessing the efficacy of AI-driven pentesting agents against complex, real-world security targets and vulnerabilities. It shifts the evaluation focus from simple task completion to validated vulnerability discovery, incorporating advanced features such as LLM-based semantic matching, continuous ground-truth maintenance, and scoring under ambiguity. Users can customize evaluations with their own targets and findings, while also accessing a set of pre-defined expert-annotated entries for standardized assessment.