ResembleAI/chatterbox-turbo

Hugging Face
AUDIO GENERATIONPricing:Input $25Concurrent Unit Cost:1Published:Dec 2, 2025License:mitArchitecture:Transformer0.7K Open Weights Warm

Resemble AI's Chatterbox-Turbo is a 350 million parameter text-to-speech (TTS) model optimized for efficient, high-quality speech generation. This model features a streamlined architecture that reduces compute and VRAM requirements, and incorporates native paralinguistic tags like [cough] and [laugh] for enhanced realism. It excels at low-latency voice agent applications, narration, and creative workflows, delivering high-fidelity audio output in a single generation step.

Loading preview...

Chatterbox-Turbo: Efficient Text-to-Speech by Resemble AI

Chatterbox-Turbo is Resemble AI's most efficient open-source text-to-speech (TTS) model, built on a streamlined 350 million parameter architecture. It significantly reduces compute and VRAM usage compared to previous models while maintaining high-fidelity audio output. A key innovation is the distillation of the speech-token-to-mel decoder, enabling speech generation in a single step, down from ten.

Key Capabilities

  • Efficient Generation: 350M parameters with reduced compute and VRAM requirements.
  • Paralinguistic Tags: Native support for tags like [cough], [laugh], and [chuckle] to add realism.
  • Single-Step Generation: Produces high-fidelity audio in one step, improving latency.
  • Zero-shot Voice Agents: Designed for real-time applications.
  • PerTh Watermarking: Includes imperceptible neural watermarks for responsible AI, surviving compression and editing.

Good For

  • Low-latency voice agents and interactive applications.
  • Narration and creative audio workflows.
  • Production environments requiring efficient and reliable TTS.
  • Zero-shot voice cloning with a reference audio clip.