ResembleAI/chatterbox

Hugging Face
AUDIO GENERATIONPricing:Input $25Concurrent Unit Cost:1Published:Apr 24, 2025License:mitArchitecture:Transformer1.8K Open Weights Warm

ResembleAI/chatterbox is a family of Text-to-Speech (TTS) models developed by Resemble AI, including the Chatterbox Multilingual V3, a 0.5 billion parameter model. It specializes in generating natural, conversational speech across 23+ languages with improved speaker similarity and reduced hallucinations. The model is designed for broad language coverage and cross-language voice cloning, offering both general multilingual and dedicated single-language finetunes.

Loading preview...

Chatterbox Multilingual V3: Advanced TTS for Global Applications

ResembleAI/chatterbox introduces Chatterbox Multilingual V3, a 0.5 billion parameter Text-to-Speech (TTS) model designed for robust multilingual speech generation. This latest iteration significantly improves speaker similarity, reduces speech hallucinations, and produces more natural, conversational output across 23+ languages. It is built on a 0.5B Llama backbone and trained on 0.5M hours of cleaned data.

Key Capabilities

  • Broad Multilingual Coverage: Supports 23+ languages including Arabic, English, Spanish, French, German, Japanese, Korean, and Chinese.
  • Enhanced Speaker Similarity: Maintains voice identity and accent preservation more consistently across different languages.
  • Reduced Hallucinations: Optimized to minimize unwanted continuations, repetitions, and off-prompt speech.
  • Exaggeration/Intensity Control: Offers unique control over speech expressiveness, allowing for more dramatic or subtle tones.
  • Single Language Pack: Provides dedicated finetunes for priority languages like Chinese, LatAm Spanish, and Hindi, offering specialized quality and dialect-aware generation.
  • Watermarked Outputs: Integrates Resemble AI's PerTh (Perceptual Threshold) Watermarker for responsible AI use.

Good For

  • Global Applications & Localization: Ideal for projects requiring high-quality speech synthesis across multiple languages.
  • Cross-Language Voice Cloning: Enables stable and reliable voice cloning where accent and identity preservation are crucial.
  • Interactive Media & AI Agents: Suitable for games, videos, and AI agents that require expressive and natural speech.
  • Dialect-Sensitive Applications: The Single Language Pack caters to use cases demanding precise language- and region-specific performance.