Vikhrmodels/salt-qwen2.5-0.5b-asr
Vikhrmodels/salt-qwen2.5-0.5b-asr is a 0.5 billion parameter model developed by Vikhrmodels, extending a pre-trained LLM with audio tokens for Automatic Speech Recognition (ASR) tasks. Utilizing SpeechTokenizer for audio tokenization, this model is specifically fine-tuned to convert spoken language into text. It achieves a Character Error Rate (CER) of 8.42 and a Word Error Rate (WER) of 18.49, making it suitable for speech-to-text applications.
Loading preview...
Model Overview
Vikhrmodels/salt-qwen2.5-0.5b-asr is a 0.5 billion parameter model designed for Automatic Speech Recognition (ASR). It extends a pre-trained Large Language Model (LLM) by integrating audio tokens and fine-tuning it specifically for speech-to-text conversion tasks. The model leverages SpeechTokenizer for efficient audio tokenization, focusing on semantic tokens.
Key Capabilities
- Automatic Speech Recognition (ASR): Converts spoken audio into text.
- Audio Tokenization: Employs SpeechTokenizer to process audio inputs.
- Performance Metrics: Achieves a Character Error Rate (CER) of 8.42 and a Word Error Rate (WER) of 18.49, indicating its accuracy in transcribing speech.
Good For
- Speech-to-Text Applications: Ideal for scenarios requiring the conversion of spoken language into written text.
- Research in ASR with LLMs: Provides a foundation for exploring how pre-trained LLMs can be adapted and enhanced for audio processing tasks.
For more technical details and code, refer to the GitHub Repo.