hexgrad/Kokoro-82M

Hugging Face
AUDIO GENERATIONPricing:Input $4Concurrent Unit Cost:1Model Size:0.082BQuant:BF16Published:Dec 26, 2024License:apache-2.0Architecture:Transformer6.8K Open Weights Warm

hexgrad/Kokoro-82M is an open-weight Text-to-Speech (TTS) model with 82 million parameters, developed by hexgrad. Based on the StyleTTS 2 and ISTFTNet architectures, it offers comparable audio quality to larger models while being significantly faster and more cost-efficient. This Apache-licensed model is optimized for flexible deployment in production environments and personal projects, providing high-quality speech synthesis at a low operational cost.

Loading preview...

hexgrad/Kokoro-82M: Lightweight and Efficient Text-to-Speech

Kokoro-82M is an open-weight Text-to-Speech (TTS) model developed by hexgrad, featuring a compact 82 million parameter architecture. Despite its small size, it achieves audio quality comparable to much larger models, offering significant advantages in speed and cost-efficiency for speech synthesis tasks. The model is Apache-licensed, enabling broad deployment across various applications.

Key Capabilities & Features

  • Lightweight Architecture: 82 million parameters, making it highly efficient for deployment.
  • High-Quality Audio: Delivers audio quality on par with larger, more resource-intensive TTS models.
  • Cost-Effective: Market rates for API usage are under $1 per million characters of text input, or approximately $0.06 per hour of audio output.
  • Open-Weight & Apache-Licensed: Provides flexibility for commercial and personal projects without restrictive licensing.
  • Architectural Foundation: Built upon the StyleTTS 2 and ISTFTNet architectures, focusing on a decoder-only design without diffusion or encoder release.
  • Permissive Training Data: Trained exclusively on permissive/non-copyrighted audio data and IPA phoneme labels, including public domain, Apache/MIT licensed audio, and synthetic audio from closed TTS models.
  • Multilingual Support: Supports multiple languages for speech generation.

When to Use This Model

  • Resource-Constrained Environments: Ideal for applications where computational resources or deployment costs are a concern.
  • Cost-Sensitive Projects: Excellent choice for projects requiring high-volume TTS at a low operational cost.
  • Flexible Deployment: Suitable for integration into production APIs, personal projects, and commercial applications due to its Apache license.
  • Rapid Prototyping: Its efficiency and ease of use (demonstrated with simple Python usage examples) make it suitable for quick development cycles.
  • Applications Requiring High-Quality, Efficient TTS: When balancing audio fidelity with performance and cost is critical.