> cat /dev/github | grep security-tools
discovered 13 Aug 2026

redteam-ai-benchmark

Python ★ 17 via github-topic
→ View on GitHub
Red Team AI Benchmark is a command-line interface tool designed to evaluate large language models (LLMs) in terms of their understanding and response quality regarding red-team-related questions and scenarios. It employs a rubric-based dataset for comprehensive assessment over 60 domain-specific questions, providing detailed metrics such as refusal rate and lexical coverage to ensure robust evaluation without executing any model outputs or engaging in any red-team activities. This tool primarily aids researchers and practitioners in assessing LLM capabilities in security contexts, with results intended for authorized use only.