Vikhrmodels/salt-qwen2.5-0.5b-tts
Vikhrmodels/salt-qwen2.5-0.5b-tts is a 0.5 billion parameter text-to-speech (TTS) model developed by Vikhrmodels, built upon the Qwen2.5 architecture. This model extends a pre-trained large language model by integrating audio tokens and fine-tuning specifically for TTS tasks. It utilizes the BigCodec tokenizer, which provides support for Slavic languages, making it suitable for speech synthesis in these linguistic contexts. The model demonstrates competitive performance in speech quality metrics like PESQ, STOI, and SI-SDR, offering a specialized solution for generating speech from text.
Loading preview...
Model Overview
Vikhrmodels/salt-qwen2.5-0.5b-tts is a 0.5 billion parameter text-to-speech (TTS) model developed by Vikhrmodels. It is built by extending a pre-trained large language model (LLM) with audio tokens and subsequently fine-tuning it for the TTS task. This approach leverages the capabilities of existing LLMs while specializing them for speech generation.
Key Features and Training
- Architecture: Based on the Qwen2.5 model family, adapted for speech synthesis.
- Tokenizer: Employs the BigCodec tokenizer, which is notable for its support of Slavic languages, enhancing its utility for specific linguistic applications.
- Training: The model underwent training for approximately 100 H100 GPU hours, indicating a focused optimization process for its TTS capabilities.
Performance Metrics
The model's performance is evaluated using standard speech quality metrics:
- PESQ (Perceptual Evaluation of Speech Quality): Achieves 1.11, indicating its perceptual quality.
- STOI (Short-Time Objective Intelligibility): Scores 0.16, reflecting its intelligibility.
- SI-SDR (Scale-Invariant Signal-to-Distortion Ratio): Reaches 23.58, demonstrating its signal reconstruction quality.
These metrics position SALT-tts competitively against similar models like Fish-aduio-1.5 and SALT-tts+asr in specific performance aspects.
Use Cases
This model is particularly well-suited for:
- Text-to-Speech applications: Generating natural-sounding speech from text inputs.
- Slavic language speech synthesis: Benefiting from the BigCodec tokenizer's support for these languages.
- Integration into applications requiring efficient speech generation: Due to its relatively compact 0.5B parameter size.