discovered 03 Aug 2026
LLMDataHub
→ View on GitHubLLMDataHub is a curated repository that collects high-quality datasets designed specifically for training large language models (LLMs). Its primary use case is to facilitate the development of chatbot models by providing comprehensive information about datasets, including their size, language, and intended usage. Notable features include a wide selection of alignment, domain-specific, and multimodal datasets, along with detailed descriptions to aid researchers and practitioners in dataset selection for improving chatbot performance.