GaborMadarasz/gemma_3_270m_ITN_finetune_16bit

TEXT GENERATIONConcurrent Unit Cost:1Model Size:0.3BQuant:BF16Context Size:32kPublished:Jul 8, 2026License:apache-2.0Architecture:Transformer Open Weights Featherless Exclusive Cold

GaborMadarasz/gemma_3_270m_ITN_finetune_16bit is a 270 million parameter Gemma-3 Instruct model, fine-tuned by Gabor Madarasz for Inverse Text Normalization (ITN) in Hungarian. It automatically converts spoken numbers, dates, and units from Automatic Speech Recognition (ASR) output into normalized written forms. This model is specifically optimized for post-processing Hungarian ASR transcripts, ensuring accurate numerical representation in text.

Loading preview...

Model Overview

This model, GaborMadarasz/gemma_3_270m_ITN_finetune_16bit, is a Gemma-3 270M Instruct model specifically fine-tuned for Inverse Text Normalization (ITN) in Hungarian. Developed by Gabor Madarasz, it aims to transform spoken numerical expressions (numbers, dates, measurements) from Automatic Speech Recognition (ASR) outputs into their standardized written forms.

Key Capabilities

  • Hungarian ITN: Converts spoken Hungarian numbers, dates, and units into normalized text (e.g., "negyven kilométer" to "40 kilométer").
  • ASR Post-processing: Designed to clean and standardize the output of Hungarian ASR systems.
  • Compact Size: Based on a 270M parameter Gemma-3 model, making it efficient for deployment.
  • High Accuracy: Achieves strong performance on dedicated test sets, with an overall Exact Match of 96.25% and Number Accuracy of 99.00% across Common Voice, VoxPopuli, and FLEURS Hungarian test splits.

Intended Use Cases

  • ASR Output Normalization: Ideal for post-processing speech recognition transcripts to ensure correct numerical formatting.
  • Automatic Subtitling: Enhances the readability and accuracy of automatically generated Hungarian subtitles.
  • Dictated Text Formatting: Useful for formatting dictated texts where numbers and units need to be standardized.

Limitations

  • Language Specific: Optimized exclusively for Hungarian; not suitable for other languages.
  • Task Specific: Not intended for general generative tasks or correcting spelling errors unrelated to ITN.
  • Aggressive Normalization: May sometimes incorrectly normalize words that contain number roots but are not actual numbers (e.g., "egészen" to "1.").
  • Domain Dependency: Training data primarily from Common Voice, VoxPopuli, and Whisper may not cover all Hungarian speech types (e.g., dialects, specialized jargon).