Vikhrmodels/salt-qwen2.5-0.5b-asr-tts

TEXT GENERATIONPricing:Input $0.04 / Cached $0.008 / Output $0.08Concurrent Unit Cost:1Model Size:0.5BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Apr 29, 2025Architecture:Transformer0.0K Featherless Exclusive Cold

Vikhrmodels/salt-qwen2.5-0.5b-asr-tts is a 0.5 billion parameter model developed by Vikhrmodels that extends a pre-trained LLM with audio tokens. It is fine-tuned for both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) tasks, utilizing a unified language model loss for dual functionality. This model is designed for efficient speech processing, offering both speech generation and recognition capabilities with minimal training overhead.

Loading preview...

Model Overview

Vikhrmodels/salt-qwen2.5-0.5b-asr-tts is a 0.5 billion parameter model developed by Vikhrmodels, designed to handle both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) tasks. It achieves this by extending a pre-trained Large Language Model (LLM) with audio tokens and fine-tuning it specifically for these dual functionalities.

Key Capabilities & Features

  • Dual-Task Functionality: Unified language model loss allows for simultaneous TTS and ASR capabilities.
  • Efficient Training: Achieves its capabilities with minimal training overhead, requiring 168 H100 GPU hours.
  • Tokenizer Support: Utilizes BigCodec tokenizer for speech generation, which supports Slavic languages, and SpeechTokenizer (semantic tokens only) for speech recognition.

Performance Metrics

The model's performance is evaluated using standard speech quality metrics:

  • PESQ (Perceptual Evaluation of Speech Quality): Achieves 1.09.
  • STOI (Short-Time Objective Intelligibility): Achieves 0.18.
  • SI-SDR (Scale-Invariant Signal-to-Distortion Ratio): Achieves 23.09.

Use Cases

This model is suitable for applications requiring integrated speech generation and recognition, particularly where efficiency and a unified approach to audio processing are beneficial. Its support for Slavic languages via the BigCodec tokenizer makes it potentially valuable for multilingual applications.