> cat /dev/github | grep security-tools
discovered 03 Aug 2026

OBLITERATUS

via awesome-list
→ View on GitHub
OBLITERATUS is an advanced open-source toolkit designed for the mechanistic interpretability of large language models, specifically focusing on the identification and removal of refusal behaviors without the need for retraining. It employs a range of techniques for abliteration, allowing researchers to visualize and manipulate internal model representations while contributing to a crowd-sourced dataset that enhances future understanding of model alignment and behavior. The tool features a user-friendly Gradio interface and a comprehensive Python API, facilitating both casual use and in-depth analysis for researchers.