discovered 03 Aug 2026
AIRTBench-Code
→ View on GitHubAIRTBench is an autonomous AI red teaming agent designed to evaluate the adversarial capabilities of large language models (LLMs) through AI/ML Capture The Flag (CTF) challenges. It systematically targets LLM-based systems to exploit vulnerabilities, providing a standardized benchmark for measuring their performance in red teaming scenarios. Notable features include a modular architecture for extensibility, integration with the Dreadnode Strikes platform, and comprehensive documentation for setup and usage.