pankaj0/llama-3.2-3b-asr-refiner-merged

TEXT GENERATIONPricing:Input $0.2036 / Output $1.34Concurrent Unit Cost:1Model Size:3.2BQuant:BF16Context Size:32kTool Calling:SupportedPublished:Oct 1, 2026License:apache-2.0Architecture:Transformer0.0K Open Weights Featherless Exclusive Cold

pankaj0/llama-3.2-3b-asr-refiner-merged is a 3.2 billion parameter language model, based on the Llama-3.1 chat template, specifically fine-tuned for speech-to-text (ASR) post-processing. It excels at correcting transcription misspellings, adding proper casing and punctuation, and removing speech disfluencies while preserving original meaning. This model is designed for refining raw ASR outputs into clean, readable text, making it ideal for applications requiring high-quality transcriptions.

Loading preview...

pankaj0/llama-3.2-3b-asr-refiner-merged: ASR Post-Processing Model

This model is a 3.2 billion parameter language model, built upon the Llama-3.1 chat template, specifically engineered for advanced speech-to-text (ASR) post-processing. It leverages a 32768 token context length to effectively refine raw ASR outputs.

Key Capabilities

  • Transcription Correction: Automatically fixes misspellings in raw ASR transcripts.
  • Punctuation and Casing: Adds correct punctuation and applies proper casing to improve readability.
  • Disfluency Removal: Eliminates common speech disfluencies such as "um," "uh," and false starts.
  • Meaning Preservation: Ensures the original meaning and exact phrasing of the speaker are maintained, avoiding summarization or rewriting.
  • Inference Optimization: Designed for efficient inference using unsloth.FastLanguageModel.

Good For

  • Enhancing ASR Quality: Ideal for applications where raw ASR output needs significant improvement in accuracy and readability.
  • Transcription Services: Useful for services requiring polished, professional-grade transcriptions.
  • Voice Assistant Backend: Can be integrated into systems that process spoken language and require clean text for further analysis or display.
  • Content Creation: Assisting in generating clean text from spoken content for articles, reports, or captions.