discovered 03 Aug 2026
exllamav2
→ View on GitHubExLlamaV2 is an inference library designed for executing local large language models (LLMs) on modern consumer GPUs, providing high-performance capabilities through features like paged attention support and a consolidated dynamic generator API. It integrates seamlessly with the TabbyAPI backend to facilitate OpenAI-compatible local or remote inference, enabling advanced functionalities such as dynamic batching and smart prompt caching. While this project is currently archived, its enhancements focus on improving inference speed and efficiency, making it suitable for various LLM applications.