sulabhkatiyar/en-indic-translate-2b
The sulabhkatiyar/en-indic-translate-2b is a 5.1 billion parameter English to 11 Indic language translation model, fine-tuned from google/gemma-4-E2B-it. It specializes in translating complex scientific documents, preserving LaTeX formulas, code blocks, and document structure across languages like Hindi, Bengali, and Tamil. This model excels at maintaining structural integrity and math preservation in long, densely structured texts, outperforming baselines in rendering and heading/code-fence matching.
Loading preview...
Overview
This model, sulabhkatiyar/en-indic-translate-2b, is a 5.1 billion parameter English to 11 Indic language translation model, fine-tuned from google/gemma-4-E2B-it. It is specifically designed for translating complex scientific documents, focusing on preserving LaTeX formulas, code blocks, and overall document structure.
Key Capabilities
- Multilingual Translation: Translates English to 11 Indic languages: Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Odia, Punjabi, Tamil, and Telugu.
- Structure Preservation: Maintains document structure, including section headings and fenced code blocks, with high integrity (0.963).
- Math Preservation: Effectively preserves LaTeX equations and symbols, achieving 0.934 math preservation and significantly reducing dropped inline math elements compared to baselines.
- Reduced Degeneracy: Exhibits a lower degeneracy (loop) rate of 7.8% and shorter loop periods, indicating more stable and usable output for long documents.
- Renderable Output: Produces translated documents that display correctly with math and markup, with a 96.2% renderable rate.
Differentiator
Unlike general-purpose translation models, en-indic-translate-2b is optimized for the unique challenges of scientific and technical document translation. It significantly outperforms the sarvamai/sarvam-translate baseline in preserving document structure, math, and code blocks, making it suitable for contexts where content integrity is paramount. The model was evaluated on a held-out set of 500 complex English documents using reference-free metrics, confirming its superior performance in maintaining the coherence and renderability of translated technical texts.