discovered 03 Aug 2026
Would-You-Kindly
→ View on GitHubWouldYouKindly is a security testing tool designed to assess the resilience of large language models (LLMs) against simulated security attacks. It features customizable attack simulations utilizing natural language prompts, a dual agent setup for realistic defense scenarios, and automated scoring to evaluate an LLM’s performance in safeguarding sensitive information. This tool is particularly useful for developers and researchers aiming to enhance LLM security and robustness.