discovered 03 Aug 2026
Speech
→ View on GitHubNVIDIA NeMo Speech is a toolkit designed for building state-of-the-art speech processing applications, including text-to-speech (TTS) and automatic speech recognition (ASR). It supports multiple languages and features advanced models such as the Nemotron series, which provide optimized performance with controllable latency, and offers capabilities for both online streaming and offline inference. Notable features include high-quality voice synthesis, low-latency conversation support, and integrations with demos hosted on Hugging Face for user testing and experimentation.