discovered 03 Aug 2026
AgentPoison
→ View on GitHubAgentPoison provides a framework for red-teaming large language model (LLM) agents through the technique of memory or knowledge base backdoor poisoning. Its primary use case is to facilitate the identification of vulnerabilities in LLMs by allowing users to optimize triggers targeting specific agent behaviors. Notable features include support for various retriever-augmented generation (RAG) embedders, configuration customization via YAML files, and trigger optimization capabilities for multiple agent types.