discovered 03 Aug 2026
VoxCPM
→ View on GitHubVoxCPM2 is a tokenizer-free Text-to-Speech (TTS) system that leverages a 2B parameter diffusion autoregressive architecture to produce high-quality, expressive speech directly from text without the need for discrete tokenization. It supports synthesis in 30 languages and offers notable features such as voice design from natural language descriptions, controllable voice cloning from short audio clips, and the output of 48kHz studio-quality audio, making it suitable for diverse multilingual speech generation tasks and creative voice applications.