ResembleAI/chatterbox-turbo
Resemble AI's Chatterbox-Turbo is a 350 million parameter text-to-speech (TTS) model optimized for efficient, high-quality speech generation. This model features a streamlined architecture that reduces compute and VRAM requirements, and incorporates native paralinguistic tags like [cough] and [laugh] for enhanced realism. It excels at low-latency voice agent applications, narration, and creative workflows, delivering high-fidelity audio output in a single generation step.
Loading preview...
Chatterbox-Turbo: Efficient Text-to-Speech by Resemble AI
Chatterbox-Turbo is Resemble AI's most efficient open-source text-to-speech (TTS) model, built on a streamlined 350 million parameter architecture. It significantly reduces compute and VRAM usage compared to previous models while maintaining high-fidelity audio output. A key innovation is the distillation of the speech-token-to-mel decoder, enabling speech generation in a single step, down from ten.
Key Capabilities
- Efficient Generation: 350M parameters with reduced compute and VRAM requirements.
- Paralinguistic Tags: Native support for tags like
[cough],[laugh], and[chuckle]to add realism. - Single-Step Generation: Produces high-fidelity audio in one step, improving latency.
- Zero-shot Voice Agents: Designed for real-time applications.
- PerTh Watermarking: Includes imperceptible neural watermarks for responsible AI, surviving compression and editing.
Good For
- Low-latency voice agents and interactive applications.
- Narration and creative audio workflows.
- Production environments requiring efficient and reliable TTS.
- Zero-shot voice cloning with a reference audio clip.