Qwen/Qwen3-TTS-12Hz-1.7B-Base
Qwen/Qwen3-TTS-12Hz-1.7B-Base is a 1.7 billion parameter base model from the Qwen3-TTS family, developed by Qwen. This model is designed for text-to-speech (TTS) applications, specifically excelling at 3-second rapid voice cloning from user audio input and serving as a base for fine-tuning other models. It supports 10 major languages and multiple dialects, featuring efficient acoustic compression and high-dimensional semantic modeling for high-fidelity speech reconstruction.
Preview unavailable
Qwen3-TTS-12Hz-1.7B-Base Overview
Qwen3-TTS-12Hz-1.7B-Base is a 1.7 billion parameter text-to-speech (TTS) model developed by Qwen, part of the Qwen3-TTS family. It is designed as a foundational model primarily for rapid voice cloning and as a base for further fine-tuning. The model supports 10 major languages (Chinese, English, Japanese, Korean, German, French, Russian, Portuguese, Spanish, and Italian) and various dialectal voice profiles.
Key Capabilities
- Rapid Voice Cloning: The model can perform 3-second rapid voice cloning from a user-provided audio input, allowing for the synthesis of new content in a cloned voice.
- Multilingual Support: It covers 10 major languages and multiple dialects, ensuring broad applicability for global users.
- Efficient Acoustic Modeling: Utilizes the self-developed Qwen3-TTS-Tokenizer-12Hz for efficient acoustic compression and high-dimensional semantic modeling, preserving paralinguistic information and acoustic environmental features.
- Universal End-to-End Architecture: Employs a discrete multi-codebook LM architecture for full-information end-to-end speech modeling, enhancing versatility and generation efficiency.
- Low-Latency Streaming Generation: Features an innovative Dual-Track hybrid streaming generation architecture, enabling end-to-end synthesis latency as low as 97ms for real-time interactive scenarios.
Use Cases
- Voice Cloning Applications: Ideal for scenarios requiring quick voice replication from short audio samples.
- Base for Fine-tuning: Serves as a robust base model for developers to fine-tune for specific custom voice or voice design applications.
- Multilingual Speech Synthesis: Suitable for generating speech in a wide array of languages with high fidelity.
- Real-time Interactive Systems: Its low-latency streaming generation makes it suitable for applications demanding immediate audio output.