discovered 03 Aug 2026
SEC-bench
→ View on GitHubSEC-bench is an automated benchmarking framework explicitly designed to evaluate Large Language Model (LLM) agents on real-world software security tasks. Its key features include automated benchmark generation from vulnerability databases, containerized reproducible vulnerability instances, and agent-oriented performance assessments on critical tasks such as vulnerability detection and patching. The tool also offers rich reporting capabilities for enhanced performance tracking and result visualization.